DETAILED ACTION
This action is responsive to the claims filed on 02/19/2026. Claims 1-25 are pending for examination.
This action is Final.
Response to Arguments
Applicant’s arguments with respect to the 35 U.S.C. 103 traversal of the claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Such claim limitation(s) is/are: “means for vocabulary generation; means for modifying a representation; means for model generation” in claim 12.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-10, 12-21, and 23-25 are rejected under 35 U.S.C. 103 as being unpatentable over Ammar et al. ((2019, June). Network-protocol-based iot device identification. In 2019 Fourth International Conference on Fog and Mobile Edge Computing (FMEC) (pp. 204-209). IEEE.), hereafter referred to as Ammar in view of Wiki et al.((2019). Bag-of-words model. Archive.org; Wikimedia Foundation, Inc. https://web.archive.org/web/20201216044033/https://en.wikipedia.org/wiki/Bag-of-words_mode), hereinafter referred to as Wiki, and in further view of Cardinal et al., (Senoussaoui, M., Cardinal, P., & Koerich, A. L. (2019). Bag-of-audio-words based on autoencoder codebook for continuous emotion prediction. arXiv preprint arXiv:1907.04928.), hereafter referred to as Cardinal and Rebuffi et al. ((2017). icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (pp. 2001-2010).), hereafter referred to as Rebuffi and Cheng et al. (US10,558,657 B1), hereafter referred to as Cheng.
Claim 1: Ammar teaches the following limitations:
instructions; and at least one processor circuit to be programmed by the instructions to at least (Ammar, p. 205–206: “presenting a new assistant designed and implemented within the Majord’Home platform.” and p. 210: “we implemented the different modules …in our IoT laboratory that uses the Majord’Home platform.” These passages describe a software assistant composed of implemented modules running on a platform, which a person of ordinary skill would understand as instructions executed by at least one processor circuit.)
obtain fingerprint data from a first network device, the first network device not previously associated with the at least one processor circuit; (Ammar, page 205, col. 1, paragraph 4, “In our work, we parse information shared by the network protocols to identify devices as this payload contain key textual description of the device itself. This device description is presented by a bag of words model.”; page 207, col. 1, section B, paragraph 2, “The Communication module manages the communication with the Majord’Home and with the IoT devices. The Majord’Home informs it when a new device is connected to the home network. This module creates a SD-LAN between itself and the device to process the identification of this device.”. Ammar’s system parses protocol payloads (DHCP, UPnP, mDNS, HTTP headers, etc.) for each new device that connects to the network and uses this information as a device “fingerprint” for identification. It is being interpreted by the examiner that the new device connected to the home network corresponds to “a first network device not previously associated with the at least one processor circuit”, and the network-protocol payload it generates corresponds to “fingerprint data”.)
associated with a first device class; (Ammar, p. 206, section 4A, paragraph 4: “The data could be labeled by the device model for a fine-granularity device identification.” Ammar also labels data by device type. This shows that the word-based representation for each device (the keys) is stored together with a device class label (model/type). It is interpreted that the keys obtained from tokenizing fingerprint data are associated with a device class, satisfying this limitation.)
update the device classes to include the first device class; (Ammar, page 206, col. 2, paragraph 6, “These vectors are used to build a database in order to to determine the nature of new connected device. Furthermore, our final data is a p × m matrix M with each column representing a word and each row representing a device. The data could be labeled by the device model for a fine-granularity device identification”, the system in Ammar continually expands its knowledge base by adding new device feature vectors and aggregating them to form or update representations of device classes.)
Wiki in the same field of Bag-of-Words model implementation, teaches the following limitations which the above fail to teach:
tokenize the fingerprint data by splitting the fingerprint data into keys (Wiki, Bag-of-words model, Definition section, “Here are two simple text documents: (1) John likes to watch movies. Mary likes movies too. (2) Mary also likes to watch football games. Based on these two text documents, a list is constructed as follows for each document: “John”,“likes”,“to”,“watch”,“movies”,“Mary”,“likes”,“movies”,“too” … “Mary”,“also”,“likes”,“to”,“watch”,“football”,“games” … Each key is the word, and each value is the number of occurrences of that word in the given text document.”, the Bag-of-Words model explicitly splits input data into individual words and treats each word as a key.)
Ammar teaches machine learning techniques for incremental classification, both using forms of exemplar or prototype based representation learning. Wiki teaches the conventional definition of a Bag-of-Words representation model which can reasonably combine with Ammar’s use of the Bag-of-Words model. Wiki teaches a Bag-of-Words model that counts the frequency of the word and applying it as a non-negative numerical value to the key. It would have been obvious to a person of ordinary skill in the art to have incorporated the Bag-of-Words exemplar representation disclosed by Wiki with the classification methods using these representations disclosed by Ammar. A motivation for the combination is to provide a way to utilize hashing, allowing for no-memory storage of dictionary values, (Wiki, section 5, “A common alternative to using dictionaries is the hashing trick, where words are mapped directly to indices with a hashing function.[5] Thus, no memory is required to store a dictionary. Hash collisions are typically dealt via freed-up memory to increase the number of hash buckets. In practice, hashing simplifies the implementation of bag-of-words models and improves scalability.”)
Cardinal in the same field of BoW incremental learning, teaches the following limitations which the above fail to teach:
decode exemplars associated with device classes based on a previously stored vocabulary, the device classes previously associated with the at least one processor circuit; (Cardinal, page 1, col. 2, paragraph 3, describe an autoencoder (AE) used to build a Bag-of-Audio-Words representation: “An AE is a type of artificial neural network that learns efficient data coding… The dimension of the AE encoded layer represents the size of the dictionary and the values of its neurons is related to the measure of the presence of the words… Using an AE to generate a BoAW has the advantage of encompassing the roles of both BoW components i.e. building a dictionary together with an embedded assignment mechanism.”; Cardinal discloses that the AE is trained on unsupervised data to learn a dictionary (vocabulary) and that the encoded layer provides a code used to represent inputs, while the decoder reconstructs the inputs from that code. In such BoW/BoAW systems, stored exemplar representations (e.g., stored codes or histograms for previously-seen classes) are interpreted/decoded through the learned dictionary. It is being interpreted by the examiner that the learned AE dictionary constitutes a stored vocabulary for previously learned classes, that exemplar representations encoded with respect to that vocabulary are decoded via the AE’s decoder, and that those exemplars are associated with particular class labels (device classes in the context of Ammar’s device-type labeling). Thus, decoding exemplar representations using the previously learned AE vocabulary is teaching upon this limitation.)
encode the exemplars based on the vocabulary and the previously stored vocabulary. (Cardinal, abstract: “The dimension of the encoded layer of the AE corresponds to the size of the dictionary” and “the output of its neurons represents the assignment metric.” This teaches that the encoded layer represents a dictionary (vocabulary) and that an input is encoded into a vector of codeword activations. Exemplar feature vectors can thus be encoded as codes relative to the vocabulary.)
Ammar teaches machine learning techniques for incremental classification, both using forms of exemplar or prototype based representation learning. Cardinal teaches that using an autoencoder-based codebook for BoW representations yields technical advantages over conventional BoW. It would have been obvious to a person of ordinary skill in the art to replace or augment Ammar’s and Wiki’s conventional BoW vocabulary with Cardinal’s Autoencoder-based BoAW codebook, in order to obtain reduced-dimensional, learned codewords and improved prediction performance, while still operating in the same BoW/BoAW feature-space. A motivation for the combination is to provide a higher performance scoring than with other baseline BoW models, (Cardinal, abstract, “Experimental results for the continuous emotion prediction task on the AVEC 2017 audio dataset have shown an improvement of the Concordance Correlation Coefficient (CCC) from 0.225 to 0.322 for arousal dimension and from 0.244 to 0.368 for valence dimension relative to the conventional BoW version implemented in a baseline system.”).
Rebuffi in the same field of incremental learning, teaches the following limitations which the above fail to teach:
at least one memory; (Rebuffi, CVPR 2017 poster, Algorithm 2: “input K // memory size,” and page 3, col. 1: “its memory requirement will be the size of the feature extraction parameters, the storage of K exemplar images …”, expressly discuss memory size and memory requirements for the algorithm. Taken together, Ammar and Rebuffi teach a computer-implemented system with executable instructions, at least one processor circuit to run those instructions, and at least one memory to store parameters and exemplar data)
generate a subsequent incremental training batch based on the vocabulary and the exemplars, (Rebuffi, page 4, col. 1, last paragraph: “First, the training set is augmented. It consists not only of the new training examples but also of the stored exemplars.” This states that each subsequent training batch is built from current samples (encoded by the feature representation vocabulary) plus the stored exmplars.)
a first exemplar of the exemplars being in the exemplars based on a proximity of a first mean of the first exemplar to a first overall mean of a corresponding one of the device classes; (Rebuffi, sec. 2.2: “To predict a label, y*, for a new image, x, it computes a prototype vector for each class observed so far, μ₁, …, μ_t, where μ_y = 1/|P_y| Σ_{p∈P_y} φ(p) is the average feature vector of all exemplars for a class y.” This teaches computing a class prototype as the mean of exemplar feature vectors, i.e., an overall mean for each device class. Rebuffi, sec. 2.2, discussing exemplar selection via herding: exemplars are chosen such that “the mean of the exemplars is close to the mean of all examples of that class” (paraphrased, with the text describing maintaining an exemplar mean that approximates the class mean). This teaches selecting exemplars so their mean stays close to the prototype μ_y, i.e., selecting a first exemplar based on proximity of its mean to the overall class mean.)
select samples from the fingerprint data and the exemplars, based on proximities of second means of corresponding ones of the samples to second overall means of corresponding ones of the updated device classes; (Rebuffi, sec. 2.2: “it computes a prototype vector for each class observed so far, μ₁, …, μ_t, where μ_y = 1/|P_y| Σ_{p∈P_y} φ(p).” This is done for all classes currently known (“observed so far”), i.e., for the updated set of classes. This teaches computing overall means (prototypes) for the updated device classes. Rebuffi, sec. 2.2, on exemplar selection: the herding procedure chooses exemplars so that their mean feature vector approximates the prototype μ_y, ensuring that “the mean of the exemplars is close to the class mean” (paraphrased from the description). This implies selecting samples from candidate data (fingerprint-based samples and existing exemplars) based on how their means relate to μ_y, i.e., selecting samples based on proximity of their means to the overall means of the updated device classes.)
train a model based on the keys as input features and based on the updated device classes; (Rebuffi, Abstract: “new classes can be added progressively.”; section 2.2: “iCaRL uses a nearest-mean-of-exemplars classification strategy. To predict a label, y ∗ , for a new image, x, it computes a prototype vector for each class observed so far, µ1, . . . , µt, where µy = 1 |Py| P p∈Py ϕ(p) is the average feature vector of all exemplars for a class y. It also computes the feature vector of the image that should be classified and assigns the class label with most similar prototype”, In Rebuffi, the model is the NME classifier (nearest-class-mean over per-class exemplar sets). The input to this classifier is a feature vector ϕ (x) (in this combination, substituted by Ammar’s Bag-of-Words “keys” as the feature vector). The updated device classes are taught by Rebuffi’s class-incremental setting (“new classes can be added progressively”), where the model is retrained/updated as additional classes arrive. Under BRI, this teaches training a model based on the keys as input features and based on the updated device classes.)
Ammar, Wiki, Cardinal and Rebuffi teach machine learning techniques for incremental classification, both using forms of exemplar or prototype based representation learning. Ammar teaches a network device identification and fingerprinting method that features using a Bag-of-Words model for determining textual representations of device data over a network; for which a conventional machine learning algorithm can be used to determine classifications, (Ammar, page 207, paragraph 1, “The textual description of a new devices is represented by a feature vector used in our data taking into account only words in the data vector. Then, a traditional method is implemented to compare the obtained feature vector with all the inputs in the database and select the most similar one.”). Rebuffi teaches a classification method by averaging all exemplars for a given class to compare to an average of a feature vector for a given device in order to determine a closest similarity. It would have been obvious to a person of ordinary skill in the art to have incorporated the Bag-of-Words exemplar representation disclosed by Ammar, Wiki, and Cardinal with the classification method using these representations disclosed by Rebuffi. A motivation for the combination is to provide a way to control the networks outputs from forgetting recent data in a catastrophic way, (Rebuffi, Page 3, col. 2, paragraph 1, “Otherwise, the network outputs will change uncontrollably, which is observable as catastrophic forgetting. In contrast, the nearest mean-of-exemplars rule (2) does not have decoupled weight vectors. The class-prototypes automatically change whenever the feature representation changes, making the classifier robust against changes of the feature representation.”).
Cheng, in the same field of text processing, topic modeling, progressive model updating, and controlled vocabulary construction, teaches the following limitation which Ammar, Wiki, Cardinal, and Rebuffi do not expressly teach:
build a vocabulary including a subset of the keys mapped to values representative of frequencies of occurrence of the keys in the fingerprint data, wherein the subset of the keys is selected based on the frequencies of occurrence of the keys satisfying a threshold and a frequency of occurrence of the keys in a previously stored vocabulary; (Cheng, col. 3, line 43, “When processing a large number of documents or a dynamic set of documents, the potential vocabulary for topic modeling can be extremely large or even unbounded, which requires a significant amount of computing and storage resources. A controlled vocabulary can be utilized in conjunction with the progressive topic modeling mechanism disclosed herein to limit the size of the vocabulary without reducing the effectiveness of the topic modeling. Because the progressive topic modeling mechanism disclosed herein processes a relatively small number of documents at each stage, the vocabulary used in a stage can be limited to a set of words found in that stage and any previous stages. For a new stage of topic modeling, the vocabulary from the previous stages can be updated to include new words from the new group of documents and to remove words that are no longer predominant in the new group of documents. The topic modeling can then be performed based on the updated vocabulary.”;
Cheng, col. 11, line 9, “As shown in FIG. 3, the vocabulary 120A used in the topic modeling for the data set 202 can be limited to a smaller size to include only words appearing in the documents 112 in the data set 202, such as “product,” “weight” “recommend,” “accurate” and “quality.” When the data set 204 is obtained to build the topic model 114B, the vocabulary 120A needs to be updated to take the new documents in the data set 204 into account. In one configuration, the update can be performed by removing from the vocabulary 120A words that appear less than a pre-defined frequency threshold, and add those that appear more than the threshold… As shown in FIG. 3, when the data set 204 is obtained or otherwise becomes available to be processed, raw word counts 302 can be calculated for the documents contained in the data set 204. Based on the raw data counts 302, a new vocabulary 120B can be generated by removing from the previous vocabulary 120A words that have a frequency less than the frequency threshold, which is set at 4 in the example shown in FIG. 3. The removed words would include “recommend” and “quality.” In the meantime, the new vocabulary 120B can also include new words appearing in the data set 204 and having a frequency higher than or equal to the threshold 4. This leads to the words “scale,” “small” and “slippery” to be added to the vocabulary 120B shown in FIG. 3.”;
Cheng, col. 11, line 42, “a weighted count can be employed to include the word counts from the previous vocabulary.”, It is being interpreted that Cheng’s words correspond to the claimed keys, Cheng’s raw word counts/weighted counts/vocabulary metrics correspond to values representative of frequencies of occurrence, Cheng’s previous vocabulary 120A corresponds to the claimed previously stored vocabulary, and Cheng’s new/updated vocabulary 120B corresponds to the claimed vocabulary including a subset of the keys. Cheng teaches selecting the subset of such words/keys for an updated vocabulary based on occurrence frequencies satisfying a predefined threshold and based on frequency/count information associated with a previously stored vocabulary.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply Cheng’s controlled-vocabulary updating technique to the bag-of-words network-device fingerprinting system of Ammar/Wiki/Rebuffi/Cardinal, in order to maintain an up-to-date and bounded vocabulary as new device fingerprint data is obtained, while reducing vocabulary dimensionality, memory usage, and processing requirements. The combination amounts to applying a known vocabulary-management technique to a known text-based bag-of-words feature representation to obtain the predictable result of selecting and maintaining a reduced set of input features for subsequent model training and classification.
Claim 2: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Rebuffi further teaches:
wherein the fingerprint data is text data, the text data corresponding to at least one of Universal Plug and Play (UPnP) data, multicast domain name system (mDNS) data, or domain name system (DNS) data. (Ammar, page 205, section III, paragraph 1, “For this work, we selected mDNS (Bonjour) and SSDP (UPnP) as they are the most used protocols by connected devices.”)
Claim 3: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Wiki further teaches:
generating a hash map of the keys and the values, the keys being the ones of the words, the values indicating the frequencies of occurrence of the fingerprint data. (Wiki, section 1,
“BoW1 ={"John":1,"likes":2,"to":1,"watch":1,"movies":2,"Mary":1,"too":1}; BoW2 = {"Mary":1,"also":1,"likes":1,"to":1,"watch":1,"football":1,"games":1}; Each key is the word, and each value is the number of occurrences of that word in the given text document.”, each key (word) is given a value representing the frequency of that word appearing in a sample, creating a hash map represented by a vector of key and value pairs.)
Claim 4: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Rebuffi, in the same field of textual feature representation, teaches the following:
wherein the one or more of the at least one processor circuit is to update the exemplars in the at least one memory to correspond to the samples to reduce monetary costs associated with data storage and data processing. (Rebuffi, page 4, section 2.4, paragraph 1, “Whenever iCaRL encounters new classes it adjusts its exemplar set. All classes are treated equally in this, i.e., when t classes have been observed so far and K is the total number of exemplars that can be stored, iCaRL will use m = K/t exemplars (up to rounding) for each class. By this it is ensured that the available memory budget of K exemplars is always used to full extent, but never exceeded.”, the exemplar set is managed and updated to conform to memory limitations, reducing monetary costs associated with storage. Since Rebuffi is capable of reducing monetary costs, it is interpreted by the examiner that it therefore meets the intended use of the limitation.)
Claim 5: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Rebuffi further teaches:
wherein the second overall means include a third overall mean and a fourth overall mean, wherein one or more of the at least one processor circuit is to further: calculate the third overall mean and the fourth overall mean, the third overall mean based on first output vectors corresponding to first samples from the fingerprint data, the fourth overall mean based on second output vectors corresponding to second samples from the exemplars, the first samples associated with the first device class, the second samples associated with a second device class from the device classes; (Rebuffi, page 3, section 2.2: “To predict a label, y ∗ , for a new image, x, it computes a prototype vector for each class observed so far, μ₁, …, μ_t, where μ_y = (1/|P_y|) Σ_{p∈P_y} φ(p) is the average feature vector of all exemplars for a class y. It also computes the feature vector of the image that should be classified …” This shows prototypes/means are averages of output vectors φ(·) (the first output vectors). And page 4, section 2.4: “one more example of the current training set is added to the exemplar set, namely the one that causes the average feature vector over all exemplars to best approximate the average feature vector over all training examples.” Thus, the third overall mean corresponds to the average of output vectors φ(x) over training examples (“first output vectors”), and the fourth overall mean corresponds to the average of output vectors φ(x) over exemplars (“second output vectors”).)
and calculate third means and fourth means of the second means, the third means corresponding to the first output vectors, (Rebuffi, page 3, section 2.2, “iCaRL uses a nearest-mean-of-exemplars classification strategy. To predict a label, y ∗ , for a new image, x, it computes a prototype vector for each class observed so far, µ1, . . . , µt, where µy = 1 |Py| P p∈Py ϕ(p) is the average feature vector of all exemplars for a class y. It also computes the feature vector of the image that should be classified and assigns the class label with most similar prototype
PNG
media_image1.png
48
251
media_image1.png
Greyscale
” ϕ is the network’s output feature vector (the claimed “second means”), thus the third means are means of the output vectors ϕ (x) taken over the first samples (training examples), i.e., “third means … corresponding to the first output vectors.).
the fourth means corresponding to the second output vectors. (Rebuffi, page 3, section 2.2, “iCaRL uses a nearest-mean-of-exemplars classification strategy. To predict a label, y ∗ , for a new image, x, it computes a prototype vector for each class observed so far, µ1, . . . , µt, where µy = 1 |Py| P p∈Py ϕ(p) is the average feature vector of all exemplars for a class y. It also computes the feature vector of the image that should be classified and assigns the class label with most similar prototype”, the fourth means are means of the output vectors ϕ (p) taken over the second samples (exemplars), i.e., “fourth means … corresponding to the second output vectors.”).
Claim 6: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 5, Rebuffi further teaches:
wherein the samples include a first ones of the samples and second ones of the samples, the first ones of the samples including a first fraction of the first samples corresponding to ones of the third means closest to the third overall mean, the second ones of the samples including a second fraction of the second samples corresponding to ones of the fourth means closest to the fourth overall mean. (Rebuffi, page 4, section 2.4, “The procedure for removing exemplars is specified in Algorithm 5. It is particularly simple: to reduce the number of exemplars from any m0 to m, one discards the exemplars pm+1, . . . , pm0 , keeping only the examples p1, . . . , pm.”, it is interpreted by the examiner that this limitation is merely suggesting that only a fraction/subset of samples from the first samples (samples gained from new devices in training) and the second samples (samples gained from the exemplar database in memory) are used for determining updated prototypes/exemplars in memory. The procedure to remove exemplars in Rebuffi includes removing exemplars from first or second samples based on their position in the exemplar set.)
Claim 7: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Ammar further teaches:
wherein one or more of the at least one processor circuit is to update the exemplars in the at least one memory to correspond to the set of samples by: matching words included in the samples to ones of the keys; (Ammar, page 206, section 4A, paragraph 2, “The textual information could be presented on the form of a word or a set of words. Each device is represented by n extracted words from protocols packets. The number of words n is different from a device to another.”, each extracted word is a key in the Bag-of-Words model.)
determining ones of the values corresponding to the words; (Ammar, page 206, section 4A, paragraph 3, “A device is represented by a feature vector of m words (features), set to 1 if the word is present in the device description and to 0 if not”, values are set for each key (word).)
and encoding the samples to a sequence of numbers based on the ones of the values. (Ammar, page 206, section 4A, paragraph 4, “the aggregation of the feature vector of devices of the same type is necessary. Besides, a set of feature vectors of several devices models is aggregated to a unique vector representing a device type”, values for each word can be aggregated into a unique vector representing a device type.)
Claim 8: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 7, Wiki, in the same field of textual feature representation, teaches the following:
wherein one or more of the at least one processor circuit is to encode the samples to prevent storing personally identifiable information. (Wiki, section 5, “A common alternative to using dictionaries is the hashing trick, where words are mapped directly to indices with a hashing function.[5] Thus, no memory is required to store a dictionary. Hash collisions are typically dealt via freed-up memory to increase the number of hash buckets. In practice, hashing simplifies the implementation of bag-of-words models and improves scalability.”, hashing can be used to encrypt personal identifiable information.)
Claim 9: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Ammar further teaches:
wherein the device classes correspond to at least one of device types, device manufacturers, or device model information. (Ammar, page 206, col. 1, paragraph 1, “In response, the UPnP-enabled devices send NOTIFY messages to advertise about their capabilities. They contain an URL pointing to the XML description of each device. This description indicates the device vendor, type, model, OS and the different services provided by the device. Such information allows us to identify the connected device type.”, each connected device using UpNp contains device type, model and manufacturer/vendor.)
Claim 10: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 1, Ammar further teaches:
wherein one or more of the at least one processor circuit is to implement the model to identify a second device class associated with second fingerprint data corresponding to a second network device, the second device class from the updated device classes. (Ammar, page 207, col. 1, paragraph 1, “These vectors are used to build a database in order to determine the nature of new connected device. Furthermore, our final data is a p × m matrix M with each column representing a word and each row representing a device. The data could be labeled by the device model for a fine-granularity device identification. Moreover, the labels could be the device type, e.g. camera, tablet, scale, etc. In this case, the aggregation of the feature vector of devices of the same type is necessary. Besides, a set of feature vectors of several devices models is aggregated to a unique vector representing a device type… After the feature extraction and feature vectors generation, the following step consists in developing our device identification method to identify new connected devices based on prior devices and be able to discriminate between them.”, new device data is input to the Bag-of-Words model for feature representation, this corresponds with second data from a second device. Feature vectors of exemplars from an exemplar set or database are compared with the feature vectors of the new device data in order to determine if that new device falls within similarity of a device class given in the exemplar set or the updated set of device classes.)
Claims 12 and 23 has limitations substantially similar to claim 1, as such a similar analysis applies.
Claims 13 and 24 has limitations substantially similar to claim 2, as such a similar analysis applies.
Claims 14 and 25 has limitations substantially similar to claim 3, as such a similar analysis applies.
Claims 15-17 has limitations substantially similar to claim 4-6 respectively, as such a similar analysis applies.
Claims 18 has limitations substantially similar to claim 7, as such a similar analysis applies.
Claim 19 has limitations substantially similar to claim 8, as such a similar analysis applies.
Claims 20-21 has limitations substantially similar to claim 9-10 respectively, as such a similar analysis applies.
Claims 11 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Ammar in view of Wiki and in further view of Cardinal, Rebuffi, Cheng and Chen et al., ((2004, April). Policy management for network-based intrusion detection and prevention. In 2004 IEEE/IFIP Network Operations and Management Symposium (IEEE Cat. No. 04CH37507) (Vol. 2, pp. 219-232). IEEE.), hereinafter referred to as Chen.
Claim 11: Ammar, Wiki, Cardinal, Rebuffi, and Cheng teaches the limitations of claim 10, Chen, in the same field of network device analysis, teaches the following which the above fails prior art fails to teach:
wherein the one or more of the at least one processor circuit is to further enforce security policies for the second network device based on the second device class. (Chen, page 221, paragraph 1, “An IPS monitors passing network traffic and performs actions on the fly. As far as this paper is concerned, an IPS maps traffic classification and a detected event to a set of actions. That is, the policies define a function f: C * E -7 R, where C is the set of traffic classification results, E the set of intrusion detection events and R the set of responses, where a response is defined as a set of actions, The device implements the mapping by way of a traffic classifier, an intrusion detection engine, a policy moderator, and a policy enforcer… The policy moderator maps the product of C and E (i.e.. C x E) to an intermediate representation”, the traffic classifier detects traffic and classifies it. Based on the classification a policy moderator will suggest a security policy based on past representations.)
Ammar, Wiki, Cardinal, Rebuffi, and Cheng teach machine learning techniques for incremental classification, both using forms of exemplar or prototype based representation learning. Chen teaches the use of enforcing security policies on networked traffic depending on its classification. It would have been obvious to a person of ordinary skill in the art to have incorporated the policy enforcement of Chen with the classification methods disclosed by Ammar, Wiki, Cardinal, Rebuffi, and Cheng. A motivation for the combination is to provide a way to automatically enforce security policies on detected traffic, (Chen, abstract, “ARM describes the mapping from intrusion types to traffic enforcement actions. It allows policies to dictate what actions to take on what types or stages of attacks. It is intuitive, and introduces a paradigm shift from flat detection rules to a structural representation that better describes an intrusion prevention system (IPS).”, combining classifications and policy management is a good advantage for Intrusion Prevention Systems, since it will provide the most accurate policy for a given trafficked activity (a device) immediately).
Claim 22 has limitations substantially similar to claim 11, as such a similar analysis applies.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Mensink, T., Verbeek, J., Perronnin, F., & Csurka, G. (2013). Distance-based image classification: Generalizing to new classes at near-zero cost. IEEE transactions on pattern analysis and machine intelligence, 35(11), 2624-2637.
Olaode, A., & Naghdy, G. (2020). Adaptive bag‐of‐visual word modelling using stacked‐autoencoder and particle swarm optimisation for the unsupervised categorisation of images. IET Image Processing, 14(9), 1769-1776.
Mai, Z., Li, R., Kim, H., & Sanner, S. (2021). Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 3589-3599).
Miettinen, M., Marchal, S., Hafeez, I., Asokan, N., Sadeghi, A. R., & Tarkoma, S. (2017, June). Iot sentinel: Automated device-type identification for security enforcement in iot. In 2017 IEEE 37th international conference on distributed computing systems (ICDCS) (pp. 2177-2184). IEEE.
Castro, F. M., Marín-Jiménez, M. J., Guil, N., Schmid, C., & Alahari, K. (2018). End-to-end incremental learning. In Proceedings of the European conference on computer vision (ECCV) (pp. 233-248).THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HYUNGJUN B YI whose telephone number is (703)756-4799. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Seed can be reached on (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/H.B.Y./Examiner, Art Unit 2146
/USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146