Prosecution Insights
Last updated: October 02, 2026
Application No. 18/437,975

DOMAIN-SPECIFIC WORD EMBEDDING MODEL

Non-Final OA §101§102§103
Filed
Feb 09, 2024
Examiner
HAN, JOSEP
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
46%
Grant Probability
Moderate
1-2
OA Rounds
1y 7m
Est. Remaining
45%
With Interview

Examiner Intelligence

Grants 46% of resolved cases
46%
Career Allowance Rate
11 granted / 24 resolved
-14.2% vs TC avg
Minimal -1% lift
Without
With
+-0.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
22 currently pending
Career history
52
Total Applications
across all art units

Statute-Specific Performance

§101
33.4%
-6.6% vs TC avg
§103
39.8%
-0.2% vs TC avg
§102
16.3%
-23.7% vs TC avg
§112
10.0%
-30.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 24 resolved cases

Office Action

§101 §102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Detailed Action The following action is in response to the communication(s) received on 02/09/2024. As of the claims filed 02/09/2024: Claims 1-20 are pending. Claims 1, 10, and 17 are independent claims. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 recites a system, thus a machine, one of the four statutory categories of patentable subject matter (Step 1). However, Claim 1 further recites: generating a plurality of content segments…, which is an evaluation or judgement that can be performed in the human mind; generating a plurality of questions based on the plurality of content segments, wherein each question of the plurality of questions corresponds to a source content segment of the plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; determining a subset of the plurality of content segments having a closest semantic similarity to at least one question of the plurality of questions, which is an evaluation or judgement that can be performed in the human mind; creating a plurality of semantic pairs, wherein each semantic pair includes the at least one question and one content segment of the subset of content segments, which is an evaluation or judgement that can be performed in the human mind; labeling a first semantic pair with a high similarity score, wherein the first semantic pair includes the source content segment corresponding to the at least one question, which is an evaluation or judgement that can be performed in the human mind; labeling at least one remaining semantic pair of the plurality of semantic pairs as a second semantic pair with a low similarity score; finetuning an embedding model using the labeled first semantic pair and the labeled second semantic pair, which is an evaluation or judgement that can be performed in the human mind; answer a network-specific user query, which is an evaluation or judgement that can be performed in the human mind. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites: comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations to finetune an embedding model for a network domain, the set of operations comprising, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application; based on a network-specific database, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application; finetuning an embedding model using the labeled first semantic pair and the labeled second semantic pair, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application; using, by a foundation model, the finetuned embedding model to retrieve context from the network-specific database, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more; implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 2, dependent on 1, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the plurality of questions is generated by a machine-learning (ML) model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 3, dependent on 2, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the ML model is one of a foundation model or a specially trained question-generation model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 4, dependent on 1, further recites determining the subset of content segments having closest semantic similarity to the at least one question is performed using a cosine similarity metric, which is a mathematical concept. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:. no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, becaus. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 5, dependent on 1, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the network-specific database is associated with a 5G multi-access edge computing system, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 6, dependent on 1, further recites the labeled first semantic pair is biased towards semantic similarity, and wherein the labeled second semantic pair is biased away from semantic similarity, which is merely a detail of an abstract idea (labeling a first semantic pair…; labeling at least one remaining semantic pair). Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:. no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, becaus. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 7, dependent on 1, further recites generating a segment vector for each content segment of the plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; generating a question vector for the at least one question, which is an evaluation or judgement that can be performed in the human mind; based on comparing the question vector to each segment vector, determining the subset of content segments having the closest semantic similarity to the at least one question, which is an evaluation or judgement that can be performed in the human mind; Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:. no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, becaus. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 8, dependent on 7, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: generating the question vector and each segment vector is performed by the embedding model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 9, dependent on 1, further recites the subset of content segments comprises top-K content segments having the closest semantic similarity to the at least one question, which is an evaluation or judgement that can be performed in the human mind. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:. no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, becaus. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 10 recites a method, thus a process, one of the four statutory categories of patentable subject matter (Step 1). However, Claim 10 further recites: generating a plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; creating a query vector representing the domain-specific query, which is an evaluation or judgement that can be performed in the human mind; creating a segment vector representing each content segment of the plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; based on comparing the query vector to each segment vector, determining a subset of the plurality of content segments having a closest semantic similarity to the domain-specific query, which is an evaluation or judgement that can be performed in the human mind; creating a prompt based on the domain-specific query and the subset of content segments, which is an evaluation or judgement that can be performed in the human mind; generating an answer to the domain-specific query based on the prompt, which is an evaluation or judgement that can be performed in the human mind. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites: using a finetuned embedding model to respond to a domain-specific query, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application; receiving a domain-specific query, which is merely an insignificant extra-solution activity of data gathering, which by MPEP 2106.05(g) cannot integrate an abstract idea into a practical application; using a finetuned embedding model,...; using the finetuned embedding model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application; and using a foundation model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the activity of data gathering (MPEP 2106.05(g)) cannot provide significantly more, as storing and retrieving information in memory is well understood, routine, and conventional (MPEP 2106.05(d)(II)(iv)); implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 11, dependent on 10, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the finetuned embedding model is trained based on pseudo-labeled semantic pairs generated based on the plurality of content segments, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 12, dependent on 10, further recites recognizes semantic similarities unique to a specialized domain associated with the domain-specific query, which is an evaluation or judgement that can be performed in the human mind. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites: the finetuned embedding model recognizes, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 13, dependent on 12, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the specialized domain is a telecommunications domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 14, dependent on 12, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the specialized domain is a 5G multi-access edge computing domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 15, dependent on 10, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: providing the answer and the domain-specific query as feedback to the embedding model, which is merely an insignificant extra-solution activity of data output, which by MPEP 2106.05(g) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the activity of data output (MPEP 2106.05(g)) cannot provide significantly more, as receiving or transmitting data over a network is well understood, routine, and conventional (MPEP 2106.05(d)(II)(i), buySAFE, Inc. v. Google, Inc). The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 16, dependent on 12, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the answer is relevant to the specialized domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 17 recites a method, thus a process, one of the four statutory categories of patentable subject matter (Step 1). However, Claim 17 further recites: generating a plurality of content segments based on a domain-specific database associated with the specialized domain, which is an evaluation or judgement that can be performed in the human mind; generating a plurality of questions based on the plurality of content segments, wherein each question of the plurality of questions corresponds to a source content segment of the plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; generating a segment vector for each content segment of the plurality of content segments, which is an evaluation or judgement that can be performed in the human mind; generating a question vector for at least one question of the plurality of questions, which is an evaluation or judgement that can be performed in the human mind; based on comparing the question vector to each segment vector, determining a subset of the plurality of content segments having a closest semantic similarity to the at least one question, which is an evaluation or judgement that can be performed in the human mind; creating a plurality of question/answer (Q/A) pairs, wherein each Q/A pair includes the at least one question and one content segment of the subset of content segments, which is an evaluation or judgement that can be performed in the human mind; labeling a first Q/A pair with a high similarity score, wherein the first Q/A pair includes the source content segment corresponding to the at least one question, which is an evaluation or judgement that can be performed in the human mind; labeling at least one remaining Q/A pair of the plurality of Q/A pairs as a second Q/A pair with a low similarity score, which is an evaluation or judgement that can be performed in the human mind. Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites: of finetuning an embedding model for a specialized domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application; based on a domain-specific database associated with the specialized domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application; using an embedding model,…; using the embedding model, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application; and finetuning the embedding model based on the labeled first Q/A pair and the labeled second Q/A pair, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more; implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 18, dependent on 17, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the specialized domain is a 5G network domain, which merely specifies the particular field of use or particular technological environment in which the abstract idea is to be performed, which by MPEP 2106.05(h) cannot integrate the abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because the particular field of use or particular technological environment (MPEP 2106.05(h)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 19, dependent on 17, further recites the labeled first Q/A pair is a positive Q/A pair, and wherein the labeled second Q/A pair is a negative Q/A pair, which is merely a detail of an abstract idea (labeling a first Q/A pair; labeling at least one remaining Q/A pair). Thus, the claim recites an abstract idea under Step 2A Prong 1.Under Step 2A Prong 2, the claim recites:. no additional elements which could integrate the abstract idea into a practical application or provide significantly more than the abstract idea itself. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, becaus. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim 20, dependent on 17, further recite. no additional abstract ideas. However: Under Step 2A Prong 2, the claim recites: the finetuned embedding model is utilized to provide context to a foundation model for returning answers to domain-specific queries, as the performance of an abstract idea on a computer is not more than instructions to 'apply it' on a computer, which by MPEP 2106.05(f) cannot integrate an abstract idea into a practical application. Thus, the claim is directed towards an abstract idea. Further, the additional element(s), alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more. The combination of these additional elements does not provide an inventive concept; thus, the claim remains ineligible. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 10-12, 15-17, 19, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Alzubi et al., "COBERT: COVID-19 question answering system using BERT " (hereinafter Alzubi). Regarding Claim 10, Alzubi teaches: A method of using a finetuned embedding model to respond to a domain-specific query, comprising: generating a plurality of content segments based on a domain-specific database; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the closed-domain corresponds to the content segments) receiving a domain-specific query; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale.) (Note: the input text corresponds to the received domain-specific query using a finetuned embedding model, creating a query vector representing the domain-specific query; (Alzubi [p.5 left ¶1] The COBERT pipeline architecture shown in Fig. 1 is made on the top of the Hugging Face transformers library. This system is classified into two different parts, such as Retriever and Reader, as shown in Figs. 2 and 3 respectively. Whenever a query is asked to the COBERT, as we can in Fig. 1 it is passed into the Retriever which selects k documents or articles from the corpus that are well on the way to contain the appropriate answer. Basically, it converts these documents into Tf-IDF vectorizer along with an input query and evaluates the cosine similarity of the document and input query.) (Note: the COBERT pipeline which contains the tf-idf vectorizer corresponds to the embedding model) using the finetuned embedding model, creating a segment vector representing each content segment of the plurality of content segments; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments; the narrowed down articles correspond to the plurality of segment vectors) based on comparing the query vector to each segment vector, determining a subset of the plurality of content segments having a closest semantic similarity to the domain-specific query; (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer.) (Note: the top 500 documents correspond to the subset of the plurality of content segments) creating a prompt based on the domain-specific query and the subset of content segments; and using a foundation model, generating an answer to the domain-specific query based on the prompt. (Alzubi [p.2 left last ¶] The reader then does the extraction of the documents by further splitting them into sentences and these sentences are then Fine-tuned with Bidirectional Encoder Representations from Transformers (BERT) for refining the automatically generated answers based on the similarity between the query. We have also used DistilBERT (i.e., Distilled-BERT) for the task of fine-tuning. Then comes our Ranker which ranks the most accurate answers using weighted scores threshold score obtained from the Retrieval and Reader. Ranker compares the scores of each candidate answer from each paragraph and compares the logits of the best answer for each paragraph. Then the best answer from that document is produced as output to the user.) Regarding Claim 11, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 10. Alzubi further teaches: The method of claim 10, wherein the finetuned embedding model is trained based on pseudo-labeled semantic pairs generated based on the plurality of content segments. (Alzubi [p.7 right ¶2] The files stored in the JSON format and processed in the form of chunks of size 10,000 one by one and the final output was generated in the form of .csv format having attributes like ‘title,’ ‘abstract,’ ‘paragraphs,’ where ‘paragraphs’ specifies the complete text of the specific research paper. The data then translated using Google Translator into English, as our main aim is to provide a QA system in the English language. Now the paragraphs column further split into sentences to optimize the training of the model. We have used the DistilBERT model for the training. DistilBERT model trained on SQuAD 1.1 using Knowledge Distillation and Bert-base-uncased-whole-word-masking-finetuned-squad as a teacher. A separate copy is prepared using json format according to the SQuAD1.1 dataset for evaluation of the model. [p.8 left ¶2] In addition, fine-tuning performed on the reader using an annotated dataset having the same format as the SQuAD dataset. The model needed two attributes, namely the ‘title’ and ‘paragraphs.’) (Note: the sentences annotated with attributes correspond to pseudo-labeling the semantic pairs) Regarding Claim 12, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 10. Alzubi further teaches: The method of claim 10, wherein the finetuned embedding model recognizes semantic similarities unique to a specialized domain associated with the domain-specific query. (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer. [p.4 right last ¶] In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments; the narrowed down articles correspond to the unique specialized domain) Regarding Claim 15, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 10. Alzubi further teaches: The method of claim 10, providing the answer and the domain-specific query as feedback to the embedding model. (Alzubi [p.2 left ¶1] The proposed system is a retriever-reader dual algorithmic approach that takes queries from the user as input and generates the most accurate responses consisting of a single line answer to the query, title of the literature, and a whole paragraph from the scientific literature [6,7,8]. [p.3 right ¶1] Then comes our Ranker which ranks the most accurate answers using weighted scores threshold score obtained from the Retrieval and Reader. Ranker compares the scores of each candidate answer from each paragraph and compares the logits of the best answer for each paragraph. Then the best answer from that document is produced as output to the user.) (Note: the user is part of the proposed system (the embedding model); thus, the user receiving the answer corresponds to providing the answer as feedback to the embedding model) Regarding Claim 16, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 12. Alzubi further teaches: The method of claim 12, wherein the answer is relevant to the specialized domain. (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer. [p.4 right last ¶] In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments; the narrowed down articles correspond to the unique specialized domain) Regarding Claim 17, Alzubi teaches: A method of finetuning an embedding model for a specialized domain, comprising: generating a plurality of content segments based on a domain-specific database associated with the specialized domain; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the closed-domain corresponds to the content segments) generating a plurality of questions based on the plurality of content segments, wherein each question of the plurality of questions corresponds to a source content segment of the plurality of content segments; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments) using an embedding model, generating a segment vector for each content segment of the plurality of content segments; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments; the narrowed down articles correspond to the plurality of segment vectors) using the embedding model, generating a question vector for at least one question of the plurality of questions; (Alzubi [p.5 left ¶1] The COBERT pipeline architecture shown in Fig. 1 is made on the top of the Hugging Face transformers library. This system is classified into two different parts, such as Retriever and Reader, as shown in Figs. 2 and 3 respectively. Whenever a query is asked to the COBERT, as we can in Fig. 1 it is passed into the Retriever which selects k documents or articles from the corpus that are well on the way to contain the appropriate answer. Basically, it converts these documents into Tf-IDF vectorizer along with an input query and evaluates the cosine similarity of the document and input query.) (Note: the documents converted using Tf-IDF vectorizer correspond to the segment vector and question vector; the selected k documents/articles from the retriever corresponds to the generated question vectors) based on comparing the question vector to each segment vector, determining a subset of the plurality of content segments having a closest semantic similarity to the at least one question; (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer.) creating a plurality of question/answer (Q/A) pairs, wherein each Q/A pair includes the at least one question and one content segment of the subset of content segments; (Alzubi [p.5 fig.2] PNG media_image2.png 486 673 media_image2.png Greyscale ) (Note: the cosine similarity scores with the matching document correspond to the plurality of Q/A pairs) labeling a first Q/A pair with a high similarity score, wherein the first Q/A pair includes the source content segment corresponding to the at least one question; (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the top answer correspond to each semantic pair with a high similarity score) labeling at least one remaining Q/A pair of the plurality of Q/A pairs as a second Q/A pair with a low similarity score; (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the 3rd highest scoring answer correspond to the second semantic pair with a low similarity score) and finetuning the embedding model based on the labeled first Q/A pair and the labeled second Q/A pair. (Alzubi [fig.3] PNG media_image3.png 507 969 media_image3.png Greyscale [p.6 left last ¶] Figure 3 shows that the reader used the most likely document, which is ranked by the retriever. To answer the question, the reader splits the document into a paragraph. Further, in the reader, a final layer is present in the architecture that matches with the help of an internal score function and the most frequent one based upon the scores is the final output. The reader is based upon Bert Processor using the ‘distilbert-base-uncased’ model. Section 4.6 explains the Distil BERT transformer model.) (Note: using the k most likely documents ranked by the retriever to generate the final output from the model corresponds to finetuning the embedded model) Regarding Claim 19, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 17. Alzubi further teaches: The method of claim 17, wherein the labeled first Q/A pair is a positive Q/A pair, and wherein the labeled second Q/A pair is a negative Q/A pair. (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: [0024] in the Specifications, a positive semantic pair corresponds to a high similarity score, and a negative semantic pair corresponds to a low similarity score; thus, the top answer correspond to the positive Q/A pair, while the 3rd highest pair, which is lower than the top answer, corresponds to the negative Q/A pair) Regarding Claim 20, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 17. Alzubi further teaches: The method of claim 17, wherein the finetuned embedding model is utilized to provide context to a foundation model for returning answers to domain-specific queries. (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale.) (Note: the retriever is specially trained to narrow down the input text for the reader; the BERT model from the reader corresponds to a foundational model) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-4, 6-9, 13 are rejected under 35 U.S.C. 103 as being unpatentable over Alzubi, in view of Holm et al., "Adopting neural language models for the telecom domain" (hereinafter Holm). Regarding Claim 1, Alzubi teaches: A system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations to… (Alzubi [p.7 right ¶3] We have used an Nvidia Tesla P100 GPU for training running on CUDA 9.2. 148 and cuDNN 7.4.1, having 16.4 GB RAM. The model trained using a pre-trained model of DistilBERT Tuning batches of size 400 in a chunk of 9000 rows at a time having a maximum sequence length of 384. Trained on GPU for 10,000 steps.) (Wang [p.9 left ¶2] Distributed computing facilitates rapid training and reasoning of the resource allocation model.) finetune an embedding model…, the set of operations comprising: (Alzubi [fig.1] PNG media_image4.png 474 997 media_image4.png Greyscale [abstract] The retriever is composed of a TF-IDF vectorizer capturing the top 500 documents with optimal scores. The reader which is pre-trained Bidirectional Encoder Representations from Transformers (BERT) on SQuAD 1.1 dev dataset built on top of the HuggingFace BERT transformers, refines the sentences from the filtered documents, which are then passed into ranker which compares the logits scores to produce a short answer, title of the paper and source article of extraction. The proposed DistilBERT version has outperformed previous pre-trained models obtaining an Exact Match(EM)/F1 score of 80.6/87.3 respectively.) (Note: the COBERT pipeline architecture comprising the retriever, reader, and ranker corresponds to the embedding model) generating a plurality of questions based on the plurality of content segments, wherein each question of the plurality of questions corresponds to a source content segment of the plurality of content segments; (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale. In other words, it can conclude that a closed-domain system [26] deals with questions under an exact domain, i.e., medicine or automotive maintenance. Table 1 shows a comparison of the ODQA and CDQA model. The domain-specific knowledge achieved by using this model fitted to a unique-domain database. The cdQA-suite built to facilitate the one who wants to model a closed-domain QA system. [p.5 table 1] PNG media_image1.png 174 1104 media_image1.png Greyscale ) (Note: the QA in the closed-domain corresponds to the content segments) (Wang [p.14 right ¶1] As a large language model, NetLM has been trained on an immense corpus of text data, enabling it to efficiently process natural language inquiries and making it ideal for customer service applications. Telecommunication companies can construct chatbots powered by NetLM to accurately interpret and respond to real-time customer queries. These NLP driven chatbots can scrutinize customer interactions, comprehend the context of their inquiries, and furnish relevant and precise responses. This dramatically alleviates the burden on customer support teams while elevating the overall customer experience. By deploying NetLM-fueled chatbots, telecommunication companies can offer continuous support to their clients, resulting in increased customer satisfaction and retention rates.) (Note: determining a subset of the plurality of content segments having a closest semantic similarity to at least one question of the plurality of questions; (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer.) creating a plurality of semantic pairs, wherein each semantic pair includes the at least one question and one content segment of the subset of content segments; (Alzubi [p.5 fig.2] PNG media_image2.png 486 673 media_image2.png Greyscale ) (Note: the cosine similarity scores with the matching document correspond to the plurality of semantic pairs) labeling a first semantic pair with a high similarity score, wherein the first semantic pair includes the source content segment corresponding to the at least one question; (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the top answer correspond to each semantic pair with a high similarity score) labeling at least one remaining semantic pair of the plurality of semantic pairs as a second semantic pair with a low similarity score; (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the 3rd highest scoring answer correspond to the second semantic pair with a low similarity score) finetuning an embedding model using the labeled first semantic pair and the labeled second semantic pair; (Alzubi [fig.3] PNG media_image3.png 507 969 media_image3.png Greyscale [p.6 left last ¶] Figure 3 shows that the reader used the most likely document, which is ranked by the retriever. To answer the question, the reader splits the document into a paragraph. Further, in the reader, a final layer is present in the architecture that matches with the help of an internal score function and the most frequent one based upon the scores is the final output. The reader is based upon Bert Processor using the ‘distilbert-base-uncased’ model. Section 4.6 explains the Distil BERT transformer model.) (Note: using the k most likely documents ranked by the retriever to generate the final output from the model corresponds to finetuning the embedded model) and using, by a foundation model, the finetuned embedding model to retrieve context from the … database to answer a… user query. (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the Q-A pair corresponds to the context and answer) Alzubi does not teach, but Holm further teaches: for a network domain (Holm [p.7 ¶1] We have named the dataset TeleQuAD, short for Telecom Question Answering Dataset. It currently contains over 4,000 question-answer pairs and the number is growing due to voluntary contributions of internal employees. Our evaluation results show that significant improvements can be achieved with a dataset of this size. [p.9 ¶4] In the case of the question-answering downstream task, further domain adaption of the telecom models was achieved by first fine-tuning on the general-domain SQuAD2.0 dataset for two epochs, followed by additional fine-tuning on the telecom-domain TeleQuAD dev set for two epochs.) generating a plurality of content segments based on a network-specific database; (Holm [p.7 ¶1] We have named the dataset TeleQuAD, short for Telecom Question Answering Dataset. It currently contains over 4,000 question-answer pairs and the number is growing due to voluntary contributions of internal employees. Our evaluation results show that significant improvements can be achieved with a dataset of this size.) (Note: the question-answer pairs correspond to the content segments based on a network-specific database) the network-specific database to answer a network-specific user query (Holm [p.10 ¶2] The potential use cases for language models within the telecom industry will grow in number if the models possess knowledge about the domain. To this end, our new benchmark, TeleQuAD, makes it possible to both adapt and evaluate models for the question answering task within the telecom domain.) Holm and Alzubi are analogous to the present invention because both are from the same field of endeavor of question-answering BERT systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset containing the telecom domain from Holm into Alzubi’s BERT question-answering system for a closed-domain system. The motivation would be to “both adapt and evaluate models for the question answering task within the telecom domain” (Holm [p.10 ¶2]). Regarding Claim 2, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi, via Alzubi/Holm, further teaches: The system of claim 1, wherein the plurality of questions is generated by a machine-learning (ML) model. (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale.) (Note: the retriever corresponds to a machine-learning model) Regarding Claim 3, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 2. Alzubi, via Alzubi/Holm, further teaches: The system of claim 2, wherein the ML model is one of a foundation model or a specially trained question-generation model. (Alzubi [p.4 right ¶2] The CDQA based COVID-19 Search engine implemented here is based on a retriever-reader dual algorithmic approach, as shown in Fig. 1. The COBERT is model inspired by the DrQA Open Domain Question Answer model, developed by Chen et al. [25]. The main challenge with Question and Answering as a Natural Language Understanding task is that QA models often fail to perform when asked to produce an answer for a question from a large input text. To address the above-discussed challenge, the COBERT model was broken into two steps. Firstly, narrow down the input text to the top articles where the answer might be present (The Retriever) using search (e.g., TF-IDF, BM25), and secondly, it out of the narrowed down input text, find the best potential answer (the reader) using a QA model [4]. For this reason, ODQA and its CDQA are considered as an approach to providing ”Machine Reading at Scale.) (Note: the retriever is specially trained to narrow down the input text and thus corresponds to a specially-trained question-generation model) Regarding Claim 4, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi, via Alzubi/Holm further teaches: The system of claim 1, wherein determining the subset of content segments having closest semantic similarity to the at least one question is performed using a cosine similarity metric. (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer.) Regarding Claim 6, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi, via Alzubi/Holm, further teaches: The system of claim 1, wherein the labeled first semantic pair is biased towards semantic similarity, and wherein the labeled second semantic pair is biased away from semantic similarity. (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the top answer correspond to each semantic pair with a high similarity score and thus biased towards semantic similarity; the 3rd highest answer is not the highest semantic similarity, thus biased away from semantic similarity) Regarding Claim 7, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi, via Alzubi/Holm further teaches: The system of claim 1, the set of operations further comprising: generating a segment vector for each content segment of the plurality of content segments; generating a question vector for the at least one question; (Alzubi [p.5 left ¶1] The COBERT pipeline architecture shown in Fig. 1 is made on the top of the Hugging Face transformers library. This system is classified into two different parts, such as Retriever and Reader, as shown in Figs. 2 and 3 respectively. Whenever a query is asked to the COBERT, as we can in Fig. 1 it is passed into the Retriever which selects k documents or articles from the corpus that are well on the way to contain the appropriate answer. Basically, it converts these documents into Tf-IDF vectorizer along with an input query and evaluates the cosine similarity of the document and input query.) (Note: the documents converted using Tf-IDF vectorizer correspond to the segment vector and question vector; the selected k documents/articles from the retriever corresponds to the generated question vectors) and based on comparing the question vector to each segment vector, determining the subset of content segments having the closest semantic similarity to the at least one question. (Alzubi [p.2 left last 2 ¶] Firstly, retrieval is performed throughout the whole corpus which is divided into paragraphs thus creating features based on tf-idf focusing on bigrams and unigrams. The embedding is then used to calculate the cosine similarity with the query and sequential comparison scores are obtained which are then used to Retrieve the top 500 documents with the optimal scores. Then, comparisons are performed batch-wise to get the document that is most probable to contain the answer.) Regarding Claim 8, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 7. Alzubi, via Alzubi/Holm, further teaches: The system of claim 7, wherein generating the question vector and each segment vector is performed by the embedding model. (Alzubi [p.5 left ¶1] The COBERT pipeline architecture shown in Fig. 1 is made on the top of the Hugging Face transformers library. This system is classified into two different parts, such as Retriever and Reader, as shown in Figs. 2 and 3 respectively. Whenever a query is asked to the COBERT, as we can in Fig. 1 it is passed into the Retriever which selects k documents or articles from the corpus that are well on the way to contain the appropriate answer. Basically, it converts these documents into Tf-IDF vectorizer along with an input query and evaluates the cosine similarity of the document and input query.) (Note: the COBERT pipeline which contains the tf-idf vectorizer corresponds to the embedding model) Regarding Claim 9, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi, via Alzubi/Holm, further teaches: The system of claim 1, wherein the subset of content segments comprises top-K content segments having the closest semantic similarity to the at least one question. (Alzubi [p.7 left ¶3] Once the retriever and reader models are used, our solution presents the top 3 answers based on a weighted score between the retriever score (based on TF-IDF cosine similarity explained in Sect. 4.2) and reader score (based on DistilBERT QA Q-A pair probability).) (Note: the top 3 answers correspond to top-k content segments) Regarding Claim 13, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 12. Alzubi does not teach, but Holm further teaches: The method of claim 12, wherein the specialized domain is a telecommunications domain. (Holm [p.10 ¶2] The potential use cases for language models within the telecom industry will grow in number if the models possess knowledge about the domain. To this end, our new benchmark, TeleQuAD, makes it possible to both adapt and evaluate models for the question answering task within the telecom domain.) Holm and Alzubi are analogous to the present invention because both are from the same field of endeavor of question-answering BERT systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset containing the telecom domain from Holm into Alzubi’s BERT question-answering system for a closed-domain system. The motivation would be to “both adapt and evaluate models for the question answering task within the telecom domain” (Holm [p.10 ¶2]). Claims 5, 14, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Alzubi, in view of Holm, further in view of Karim et al., "SPEC5G: A Dataset for 5G Cellular Network Protocol Analysis" (hereinafter Karim). Regarding Claim 5, Alzubi/Holm respectively teaches and incorporates the claimed limitations and rejections of Claim 1. Alzubi/Holm does not teach, but Karim further teaches: The system of claim 1, wherein the network-specific database is associated with a 5G multi-access edge computing system. (Karim [p.1 right ¶2] In this paper, we address this need by introducing SPEC5G, a high-quality dataset of the 5G protocol specifications. 5G is not a single wireless technology, but an umbrella term used to categorize the fifth generation of wireless communication, including hundreds of different protocols at different layers of the protocol. Some of these protocols are VoWiFi, cellular IoT, IKE, and 5G-AKA. SPEC5G is a complete dataset that covers all these protocols and therefore, has the potential to impact different protocols affecting billions of devices. Such a high-quality dataset would be beneficial to numerous applications in different domains, such as security testing, policy enforcement, automatic code generation, and protocol summarization. It would encourage research and development in novel NLP tasks that are communication protocol-specific and critical for the security analysis of these protocols.) Karim and Alzubi/Holm are analogous to the present invention because both are from the same field of endeavor of fine-tuning telecom-based BERT models. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset including 5g protocol specifications from Karim into Alzubi/Holm’s telecom-based BERT model. The motivation would be to “With the wealth of information present in the dataset, question-answering datasets could also be constructed, where models are trained to answer questions related to 5G concepts, protocols, and technologies.” (Karim [p.9 right ¶1]). Regarding Claim 14, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 12. Alzubi does not teach, but Holm further teaches: The method of claim 12, wherein the specialized domain is a… multi-access edge computing domain. (Holm [p.4 ¶1] In the telecom domain, there are technical terminologies, named entities, and relations among them. For example, eNB is a type of base station, user equipment (UE) is almost the same as a mobile device, and 4G and LTE refer to the same technology. Our goal has been to make a language model that can comprehend sentences where these telecom terms are used.) (Note: the telecom domain which includes UE corresponds to a multi-access edge computing domain) Holm and Alzubi are analogous to the present invention because both are from the same field of endeavor of question-answering BERT systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset containing the telecom domain from Holm into Alzubi’s BERT question-answering system for a closed-domain system. The motivation would be to “both adapt and evaluate models for the question answering task within the telecom domain” (Holm [p.10 ¶2]). Alzubi/Holm does not teach that the telecom domain includes 5G, but Karim further teaches: 5G multi-access edge computing domain (Karim [p.1 right ¶2] In this paper, we address this need by introducing SPEC5G, a high-quality dataset of the 5G protocol specifications. 5G is not a single wireless technology, but an umbrella term used to categorize the fifth generation of wireless communication, including hundreds of different protocols at different layers of the protocol. Some of these protocols are VoWiFi, cellular IoT, IKE, and 5G-AKA. SPEC5G is a complete dataset that covers all these protocols and therefore, has the potential to impact different protocols affecting billions of devices. Such a high-quality dataset would be beneficial to numerous applications in different domains, such as security testing, policy enforcement, automatic code generation, and protocol summarization. It would encourage research and development in novel NLP tasks that are communication protocol-specific and critical for the security analysis of these protocols.) Karim and Alzubi/Holm are analogous to the present invention because both are from the same field of endeavor of fine-tuning telecom-based BERT models. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset including 5g protocol specifications from Karim into Alzubi/Holm’s telecom-based BERT model. The motivation would be to “With the wealth of information present in the dataset, question-answering datasets could also be constructed, where models are trained to answer questions related to 5G concepts, protocols, and technologies.” (Karim [p.9 right ¶1]). Regarding Claim 18, Alzubi respectively teaches and incorporates the claimed limitations and rejections of Claim 17. Alzubi does not teach, but Holm further teaches: The method of claim 17, wherein the specialized domain is a 5G network domain. (Holm [p.4 ¶1] In the telecom domain, there are technical terminologies, named entities, and relations among them. For example, eNB is a type of base station, user equipment (UE) is almost the same as a mobile device, and 4G and LTE refer to the same technology. Our goal has been to make a language model that can comprehend sentences where these telecom terms are used.) (Note: the telecom domain which includes UE corresponds to a network domain) Holm and Alzubi are analogous to the present invention because both are from the same field of endeavor of question-answering BERT systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset containing the telecom domain from Holm into Alzubi’s BERT question-answering system for a closed-domain system. The motivation would be to “both adapt and evaluate models for the question answering task within the telecom domain” (Holm [p.10 ¶2]). Alzubi/Holm does not teach that the specialized telecom domain includes 5G, but Karim further teaches: 5G network domain (Karim [p.1 right ¶2] In this paper, we address this need by introducing SPEC5G, a high-quality dataset of the 5G protocol specifications. 5G is not a single wireless technology, but an umbrella term used to categorize the fifth generation of wireless communication, including hundreds of different protocols at different layers of the protocol. Some of these protocols are VoWiFi, cellular IoT, IKE, and 5G-AKA. SPEC5G is a complete dataset that covers all these protocols and therefore, has the potential to impact different protocols affecting billions of devices. Such a high-quality dataset would be beneficial to numerous applications in different domains, such as security testing, policy enforcement, automatic code generation, and protocol summarization. It would encourage research and development in novel NLP tasks that are communication protocol-specific and critical for the security analysis of these protocols.) Karim and Alzubi/Holm are analogous to the present invention because both are from the same field of endeavor of fine-tuning telecom-based BERT models. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement the dataset including 5g protocol specifications from Karim into Alzubi/Holm’s telecom-based BERT model. The motivation would be to “With the wealth of information present in the dataset, question-answering datasets could also be constructed, where models are trained to answer questions related to 5G concepts, protocols, and technologies.” (Karim [p.9 right ¶1]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEP HAN whose telephone number is (703)756-1346. The examiner can normally be reached Mon-Fri 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kakali Chaki can be reached on (571) 272-3719. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.H./Examiner, Art Unit 2122 /KAKALI CHAKI/Supervisory Patent Examiner, Art Unit 2122
Read full office action

Prosecution Timeline

Feb 09, 2024
Application Filed
Sep 11, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737435
DEEP NEURAL NETWORKS VIA PROTOTYPE FACTORIZATION
4y 10m to grant Granted Sep 15, 2026
Patent 12718087
DEEP LEARNING SOFTWARE MODEL MODIFICATION
5y 0m to grant Granted Aug 25, 2026
Patent 12705465
METHOD FOR OPERATING NEURAL NETWORK
4y 3m to grant Granted Aug 11, 2026
Patent 12688425
PREDICTION OF CLASSIFICATION OF AN UNKNOWN INPUT BY A TRAINED NEURAL NETWORK
4y 11m to grant Granted Jul 21, 2026
Patent 12651042
SYSTEM AND METHOD FOR MACHINE LEARNING FAIRNESS TESTING
4y 8m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
46%
Grant Probability
45%
With Interview (-0.7%)
4y 3m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 24 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month