Prosecution Insights
Last updated: October 02, 2026
Application No. 17/677,009

COMPRESSING INFORMATION IN AN END NODE USING AN AUTOENCODER NEURAL NETWORK

Final Rejection §103
Filed
Feb 22, 2022
Priority
Sep 17, 2020 — divisional of 11/802,894
Examiner
LEY, SALLY THI
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
Silicon Laboratories Inc.
OA Round
4 (Final)
23%
Grant Probability
At Risk
5-6
OA Rounds
2m
Est. Remaining
37%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
10 granted / 44 resolved
-32.3% vs TC avg
Moderate +15% lift
Without
With
+14.6%
Interview Lift
resolved cases with interview
Typical timeline
4y 9m
Avg Prosecution
17 currently pending
Career history
80
Total Applications
across all art units

Statute-Specific Performance

§101
24.5%
-15.5% vs TC avg
§103
55.3%
+15.3% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
9.6%
-30.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 44 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 30 Oct 2025 has been entered. Status of Claims This Office Action is in response to the communication filed 22 May 2026. Claims 1-6, 8-10, 12-13, and 17-23 are being considered on the merits. Claim Rejections – 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-6, 8, 12-14, and 17-23 are rejected under 35 U.S.C. 103 as being unpatentable over Kamath, Ajith M. (US 2014/0185862 A1; hereinafter “Kamath”) in view of Song, et. al. (US 2020/0043241 A1; hereinafter, “Song”) and in view of P. Agarwal, S. Poddar, A. Hazarika and H. Rahaman ("Learning to synthesize faces using voice clips for Cross-Modal biometric matching," 2019 IEEE Region 10 Symposium (TENSYMP), Kolkata, India, 2019, pp. 397-402; hereinafter, “Agarwal”) and further in view of Vedal, Amund Hansen (“Unsupervised Audio Spectrogram Compression using Vector Quantized Autoencoders” DEGREE PROJECT COMPUTER SCIENCE AND ENGINEERING, SECOND CYCLE, 30 CREDITS , STOCKHOLM SWEDEN 2019; hereinafter, “Vedal”). Claim 1, Kamath as modified teaches: At least one non-transitory computer readable storage medium having stored thereon instructions, which if performed by a machine cause the machine to perform a method comprising: (Kamath, para. 0081: “The methods and processes described above may be implemented in programs executed from a system's memory (a computer readable medium, such as an electronic, optical or magnetic storage device). The methods, instructions and circuitry operate on electronic signals, or signals in other electromagnetic forms.”) generating, in at least one cloud server (Kamath para 0077) comprising the machine, an autoencoder comprising an encoder and a decoder, and (Song, para. 0292: “More specifically, FIG. 9(a) illustrates a general structure of the artificial neural network model, and FIG. 9(b) illustrates an autoencoder, that performs decoding after encoding and goes through a reconstruction step, among the artificial neural network model.”) generating a classifier, (Song, para. 0276: “Furthermore, the memory 25 may store the neural network model (e.g. the deep learning model 26) generated through a learning algorithm for classifying/recognizing data in accordance with the embodiment of the present disclosure.”) wherein the encoder is to encode a spectrogram into a compressed spectrogram, (Vedal, pg. 51 and figure. 5.5: “The results for audio spectrogram compression presented in Figure 5.5 sug gests VQVAE-modelssuch as(K=1024,D=256)performsimilarlytoAE(D=3), with a 9.5x higher compression rate achieved by introducing the vector quan tization bottleneck”) the decoder is to decode the compressed spectrogram into a recovered spectrogram, and (Kamath, para. 0031: “The metadata may be encoded in the machine readable information, embedded within the audio signal, such that only intended recipients can decode it”) the classifier is to identify one or more properties of real world information comprising speech information of a user (Agarwal, sec. I: “We compare the performance of different generative models and show that adding pixel loss as a regularization factor to the original adversarial loss of conditional GANs helps in improving the accuracy of the synthesized images corresponding to the identity of the individual as specified by the voice embeddings.” Agarwal teaches identifying an identity of a speaker i.e. real world information comprising voice clip i.e. speech information of the individual) from the recovered spectrogram; (Agarwal, sec. 2: “ The reconstruction loss is represented by the first term in (2) where the expectation of the datapoint i is taken with respect to the distribution of the encoder which helps the decoder in efficient reconstruction.”) calculating a first loss of the and independently calculating a second loss of the classifier; (Agarwal, sec. IV(C): “This is mainly because RC-GAN penalizes the network for both the adversarial loss as well as the pixel loss in comparison to only adversarial loss in case of C-GAN.” Agarwal teaches a first adversarial loss as well as a second pixel loss). jointly training the autoencoder and the classifier to minimize a loss function based at least in part on the first loss and the second loss (Agarwal, sec. II and III: “The loss function for conditional GANs are given as [equation omitted] where the discriminator tries to maximize L(G, D) while the generator tries to minimize it.” “We train a voice classification network which takes as input spectrograms of dimension 60* 41* 2 as described in Section 4-A. This network is 6 layered deep as shown in the Fig. 1 and is trained using Adam Optimizer with a learning rate of 0.001 to optimize the categorical cross entropy loss.” Examiner notes Agarwal teaches minimizing an overall loss based on a first loss by a generator and a second loss by a discriminator of the GAN). storing the trained autoencoder and the trained classifier in a non-transitory storage medium for delivery of at least the encoder of the trained autoencoder to at least one end node wireless device to cause the encoder of the trained autoencoder to execute on the at least one end node wireless device to compress the real world information sensed by the at least one end node wireless device into the compressed spectrogram (Agarwal sec. IV(A): “All the audio clips are collected from Youtube videos in a challenging environment conditions. These audio clips are mp3 compressed and are converted to spectrograms as shown in Figures 5 using the same approach as in [25]”), the compressed spectrogram comprising the compressed real world information to be sent from the at least one end node wireless device to the at least one cloud server for processing; (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) receiving, in the at least one cloud server, the compressed spectrogram comprising (Agarwal sec. IV(A) as set forth above) the compressed real world information; (Kamath, para. 0055: “ In particular, as depicted in FIG. 6 for example, one user authors a message (e.g., an image) on her phone with a mobile application program, and the mobile application writes it into the spectrogram of a host audio signal, as shown in block 112. The application converts the spectrogram to an output audio signal format (e.g., way file) in block 114 and plays that audio signal.” Examiner notes Kamath para. 0077 above teaches that any functionality, including receiving can be implemented on different devices including, “e.g. a remote server”). processing, in the at least one cloud server, the compressed real world information to determine, based at least in part on the compressed real world information, an operation requested by the user to be performed by the at least one end node wireless device; and (Song, para. 0079: “At this time, the AI server 16 may receive input data from the AI device (11 to 15), infer a result value from the received input data by using the learning model, generate a response or control command based on the inferred result value, and transmit the generated response or control command to the AI device (11 to 15).” Examiner notes Song teaches generating an inferred result i.e. processing. Examiner further notes Song teaches any devices being a wireless device in parap 0138 and including a robot as taught in para. 0090) sending, from the at least one cloud server, command information to the at least one end node wireless device to cause the at least one end node wireless device to perform the operation responsive to the command information. (Song, para. 0090: “ Also, the robot 11 may perform the operation or navigate the space by controlling its locomotion platform based on the control/interaction of the user. At this time, the robot 11 may obtain intention information of the interaction due to the user's motion or voice command and perform an operation by determining a response based on the obtained intention information.” Examiner notes Song teaches an end node being a robot where Kamath teaches an end node being a cell phone. Moreover Song teaches devices being wireless at para 0138) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as modified. Kamath teaches communicating a message between devices by passing an audio signal with the message written into the spectrogram of the audio signal. Song teaches an intelligent device and a method of controlling the same. One of ordinary skill would have been motivated to combine the teachings of Song into Kamath, as modified, in order increase reliability and latency to support smart grid control, industry automation, robotics, drone control and coordination (Song, para. 0062). Additionally, it would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Agarwal into Kamath, as modified. Agarwal teaches a framework for cross-modal biometric where faces of an individual are generated using his/her voice clips and further the synthesized faces are tested using a face classification network. One of ordinary skill would have been motivated to combine the teachings of Agarwal into Kamath, as modified, in order to enable cross-domain biometric matching using generative networks (Agarwal, sec. I). Additionally, it would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vedel into Kamath, as modified. Vedel teaches quantitative performance comparisons for AE and VQVAE models on a spectrogram compression task. One of ordinary skill would have been motivated to combine the teachings of Vedel into Kamath, as modified, in order to extract observable audio features (Vedal, pg. 11). Claim 4, Kamath as modified teaches claim 1. Kamath as modified further teaches: wherein the method further comprises sending a trained encoder portion of the autoencoder to one or more end node wireless devices to enable the one or more end node wireless devices to compress spectrograms using the trained encoder portion and (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) send the compressed spectrograms to the at least one cloud server. (Kamath, para. 0049: “After writing the message into the spectrogram, the spectrogram is converted to an audio signal suitable for play out, transmission or storage (converted to a standard audio signal and file format, possibly compressed to reduce its size).” Examiner notes that Kamath teaches uploading and downloading operations between services and devices at para 0077) Claim 5, Kamath as modified teaches claim 4. Kamath as modified further teaches: wherein the method further comprises: receiving, in the at least one cloud server, uncompressed spectrograms from the one or more end node wireless devices; and (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.” Examiner notes Kamath teaches encoding and decoding signals at para. 0031 and compression of signals at 0049). incrementally training, in the at least one cloud server, one or more of the autoencoder or the classifier based at least in part on the uncompressed spectrograms. (Song, para. 0297 and 0298: “Referring to FIG. 9(b), an artificial neural network model according to an embodiment of the present disclosure may include an autoencoder. “ “The artificial neural network model that has been learned repeatedly several times may stop the learning and may be stored in a memory of an AI device if an error value is less than a reference value.”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath as set forth above with respect to claim 1. Claim 6, Kamath as modified teaches claim 5. Kamath as modified further teaches: wherein the method further comprises sending an incrementally trained encoder portion of the autoencoder from the at least one cloud server to the one or more end node wireless devices. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.” Examiner notes Kamath teaches encoding and decoding signals at para. 0031 and compression of signals at 0049). Claim 8, Kamath as modified teaches: A method comprising: generating, in at least one cloud server comprising the machine, an autoencoder comprising an encoder and a decoder, and (Song, para. 0292: “More specifically, FIG. 9(a) illustrates a general structure of the artificial neural network model, and FIG. 9(b) illustrates an autoencoder, that performs decoding after encoding and goes through a reconstruction step, among the artificial neural network model.”) generating a classifier, (Song, para. 0276: “Furthermore, the memory 25 may store the neural network model (e.g. the deep learning model 26) generated through a learning algorithm for classifying/recognizing data in accordance with the embodiment of the present disclosure.”) wherein the encoder is to encode a spectrogram into a compressed spectrogram, (Vedal, pg. 51 and figure. 5.5: “The results for audio spectrogram compression presented in Figure 5.5 sug gests VQVAE-modelssuch as(K=1024,D=256)performsimilarlytoAE(D=3), with a 9.5x higher compression rate achieved by introducing the vector quan tization bottleneck”) the decoder is to decode the compressed spectrogram into a recovered spectrogram, and (Kamath, para. 0031: “The metadata may be encoded in the machine readable information, embedded within the audio signal, such that only intended recipients can decode it”) the classifier is to identify one or more properties of real world information from the recovered spectrogram; (Song, para. 0008: “In an aspect, a method of controlling an intelligent device includes obtaining sound information from a photographed image; learning the obtained sound information and recognizing a sound based on the result of the learned sound information; and classifying the image based on the recognized sound.”) calculating a first loss of the and independently calculating a second loss of the classifier; (Agarwal, sec. IV(C): “This is mainly because RC-GAN penalizes the network for both the adversarial loss as well as the pixel loss in comparison to only adversarial loss in case of C-GAN.” Agarwal teaches a first adversarial loss as well as a second pixel loss). jointly training the autoencoder and the classifier to minimize a loss function based at least in part on the first loss and the second loss (Agarwal, sec. II and III: “The loss function for conditional GANs are given as [equation omitted] where the discriminator tries to maximize L(G, D) while the generator tries to minimize it.” “We train a voice classification network which takes as input spectrograms of dimension 60* 41* 2 as described in Section 4-A. This network is 6 layered deep as shown in the Fig. 1 and is trained using Adam Optimizer with a learning rate of 0.001 to optimize the categorical cross entropy loss.” Examiner notes Agarwal teaches minimizing an overall loss based on a first loss by a generator and a second loss by a discriminator of the GAN). storing the trained autoencoder and the trained classifier in a non-transitory storage medium for delivery of at least the encoder of the trained autoencoder to at least one end node wireless device to cause the encoder of the trained autoencoder to execute on the at least one end node wireless device to compress the real world information sensed by the at least one end node wireless device into the compressed spectrogram (Agarwal sec. IV(A): “All the audio clips are collected from Youtube videos in a challenging environment conditions. These audio clips are mp3 compressed and are converted to spectrograms as shown in Figures 5 using the same approach as in [25]”), the compressed spectrogram comprising the compressed real world information to be sent from the at least one end node wireless device to the at least one cloud server for processing; (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) receiving, in the computer system, the compressed spectrogram comprising (Agarwal sec. IV(A) as set forth above) image information; (Kamath, para. 0029: “Whether embedded or linked to the audio signal, the control data may include data for controlling distribution or access to a message. For example, the control data may include a key or pointer to a key used to descramble a spectrogram image so that it may be viewable only by authorized recipients.”) processing, in the computer system, the compressed real world information (Agarwal sec. IV(A) set forth above) comprising the image information (Kamath, para. 0029) to identify presence of a person in a local environment of the at least one end node wireless device; and (Kamath, para. 0041: “First, the presence of the user is detected and authenticated by any number of communication channels between the user's mobile device and the venue, including for example, a private audio channel, low power BlueTooth signal, or wi-fi signal, to name a few options.”) in response to identifying the presence of a person, sending from the computers system, command information to the at least one end node wireless device to cause the at least one end node wireless device to perform an operation responsive to the command information. (Kamath, para. 0077: “Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.)”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as set forth above with respect to claim 1. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Agarwal into Kamath, as modified, as set forth above with respect to claim 1. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vedel into Kamath, as modified, as set forth above with respect to claim 1. Claim 12, Kamath as modified teaches claim 8. Kamath as modified further teaches: further comprising sending a trained encoder portion of the autoencoder from the computer system to one or more end node wireless devices. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) Claim 13, Kamath as modified teaches claim 12. Kamath as modified further teaches: further sending a trained decoder portion of the autoencoder from the computer system to at least some of the one or more end node wireless devices. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) Claim 14, Kamath as modified teaches claim 12. Kamath as modified further teaches: further comprising receiving, in the computer system, compressed spectrograms from at least some of the one or more end node wireless devices, the compressed spectrograms compressed using the trained encoder portion. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) Claim 17, Kamath as modified teaches claim 12. Kamath as modified further teaches: further comprising: requesting, by the computer system, one or more uncompressed spectrograms from at least some of the one or more end node wireless devices; and (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) incrementally training, in the computer system, at least one of the autoencoder or the classifier based at least in part on the one or more uncompressed spectrograms. (Song, para. 0297 and 0298: “Referring to FIG. 9(b), an artificial neural network model according to an embodiment of the present disclosure may include an autoencoder. “ “The artificial neural network model that has been learned repeatedly several times may stop the learning and may be stored in a memory of an AI device if an error value is less than a reference value.”) Claim 18, Kamath as modified teaches: A system comprising: at least one processor; memory coupled to the at least one processor; and one or more non-transitory storage media, wherein the one or more non-transitory storage media comprises instructions which if performed by the system cause the system to perform a method comprising (Kamath, para. 0081: “The methods and processes described above may be implemented in programs executed from a system's memory (a computer readable medium, such as an electronic, optical or magnetic storage device). The methods, instructions and circuitry operate on electronic signals, or signals in other electromagnetic forms.”) generating, in at least one cloud server comprising the machine, an autoencoder comprising an encoder and a decoder, and (Song, para. 0292: “More specifically, FIG. 9(a) illustrates a general structure of the artificial neural network model, and FIG. 9(b) illustrates an autoencoder, that performs decoding after encoding and goes through a reconstruction step, among the artificial neural network model.”) generating a classifier, (Song, para. 0276: “Furthermore, the memory 25 may store the neural network model (e.g. the deep learning model 26) generated through a learning algorithm for classifying/recognizing data in accordance with the embodiment of the present disclosure.”) wherein the encoder is to encode a spectrogram into a compressed spectrogram, (Vedal, pg. 51 and figure. 5.5: “The results for audio spectrogram compression presented in Figure 5.5 sug gests VQVAE-modelssuch as(K=1024,D=256)performsimilarlytoAE(D=3), with a 9.5x higher compression rate achieved by introducing the vector quan tization bottleneck”) the decoder is to decode the compressed spectrogram into a recovered spectrogram, and (Kamath, para. 0031: “The metadata may be encoded in the machine readable information, embedded within the audio signal, such that only intended recipients can decode it”) the classifier is to identify one or more properties of real world information from the recovered spectrogram; (Agarwal, sec. I: “We compare the performance of different generative models and show that adding pixel loss as a regularization factor to the original adversarial loss of conditional GANs helps in improving the accuracy of the synthesized images corresponding to the identity of the individual as specified by the voice embeddings.” Agarwal teaches identifying an identity of a speaker i.e. real world information comprising voice clip i.e. speech information of the individual) calculating a first loss of the and independently calculating a second loss of the classifier; (Agarwal, sec. IV(C): “This is mainly because RC-GAN penalizes the network for both the adversarial loss as well as the pixel loss in comparison to only adversarial loss in case of C-GAN.” Agarwal teaches a first adversarial loss as well as a second pixel loss). jointly training the autoencoder and the classifier to minimize a loss function based at least in part on the first loss and the second loss (Agarwal, sec. II and III: “The loss function for conditional GANs are given as [equation omitted] where the discriminator tries to maximize L(G, D) while the generator tries to minimize it.” “We train a voice classification network which takes as input spectrograms of dimension 60* 41* 2 as described in Section 4-A. This network is 6 layered deep as shown in the Fig. 1 and is trained using Adam Optimizer with a learning rate of 0.001 to optimize the categorical cross entropy loss.” Examiner notes Agarwal teaches minimizing an overall loss based on a first loss by a generator and a second loss by a discriminator of the GAN). storing the trained autoencoder and the trained classifier in a non-transitory storage medium for delivery of at least the encoder of the trained autoencoder to at least one end node wireless device to cause the encoder of the trained autoencoder to execute on the at least one end node wireless device to compress the real world information sensed by the at least one end node wireless device into the compressed spectrogram (Agarwal sec. IV(A): “All the audio clips are collected from Youtube videos in a challenging environment conditions. These audio clips are mp3 compressed and are converted to spectrograms as shown in Figures 5 using the same approach as in [25]”), the compressed spectrogram comprising the compressed real world information to be sent from the at least one end node wireless device to the at least one cloud server for processing; (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) receiving, in the system, the compressed spectrogram comprising (Agarwal sec. IV(A) as set forth above) the compressed real world information; (Kamath, para. 0055: “ In particular, as depicted in FIG. 6 for example, one user authors a message (e.g., an image) on her phone with a mobile application program, and the mobile application writes it into the spectrogram of a host audio signal, as shown in block 112. The application converts the spectrogram to an output audio signal format (e.g., way file) in block 114 and plays that audio signal.” Examiner notes Kamath para. 0077 above teaches that any functionality, including receiving can be implemented on different devices including, “e.g. a remote server”). processing, in the system, the compressed real world information to determine, based at least in part on the compressed real world information, an operation requested by the user to be performed by the at least one end node wireless device; and (Song, para. 0079: “At this time, the AI server 16 may receive input data from the AI device (11 to 15), infer a result value from the received input data by using the learning model, generate a response or control command based on the inferred result value, and transmit the generated response or control command to the AI device (11 to 15).” Examiner notes Song teaches generating an inferred result i.e. processing. Examiner further notes Song teaches any devices being a wireless device in para. 0138 and including a robot as taught in para. 0090) sending, from the system, command information to the at least one end node wireless device to cause the at least one end node wireless device to perform the operation responsive to the command information. (Song, para. 0090: “ Also, the robot 11 may perform the operation or navigate the space by controlling its locomotion platform based on the control/interaction of the user. At this time, the robot 11 may obtain intention information of the interaction due to the user's motion or voice command and perform an operation by determining a response based on the obtained intention information.” Examiner notes Song teaches an end node being a robot where Kamath teaches an end node being a cell phone. Moreover Song teaches devices being wireless at para 0138) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as set forth above with respect to claim 1. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Agarwal into Kamath, as modified, as set forth above with respect to claim 1. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Vedel into Kamath, as modified, as set forth above with respect to claim 1. Claim 19, Kamath as modified teaches claim 18. Kamath as modified further teaches: wherein the system comprises a remote cloud server, the remote cloud server to send the trained autoencoder (Song para 0292) to one or more end nodes coupled to the remote cloud server via a network. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as set forth above with respect to claim 1. Claim 20, Kamath as modified teaches claim 19. Kamath as modified further teaches: wherein the remote cloud server is to receive compressed spectrograms (Kamath, para. 0049: “After writing the message into the spectrogram, the spectrogram is converted to an audio signal suitable for play out, transmission or storage (converted to a standard audio signal and file format, possibly compressed to reduce its size).”) from at least some of the one or more end node devices and process the compressed spectrograms. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) Claim 21, Kamath as modified teaches claim 1. Kamath as modified further teaches: wherein the command information is to cause the at least one end node wireless device to perform the operation comprising a playing of a media file. (Kamath, para. 0077: “Different of the functionality can be implemented on different devices. For example, in a system in which a cell phone communicates with a server at a remote service provider, different tasks can be performed exclusively by one device or the other, or execution can be distributed between the devices. For example, messages can be authored and communicated to other devices by servers in a cloud computing service by uploading message and host image content to servers in a cloud service or authored in a mobile device via a script program downloaded from an online authoring service provided from a network server. Also, messages and host signals may be stored on the cell phone—allowing the cell phone to write messages into host signals, transmit them, receive them, and render them—all without reliance on externals devices. Thus, it should be understood that description of an operation as being performed by a particular device (e.g., a cell phone) is not limiting but exemplary; performance of the operation by another device (e.g., a remote server), or shared between devices, is also expressly contemplated. (Moreover, more than two devices may commonly be employed. E.g., a service provider may refer some tasks, functions or operations, to servers dedicated to such tasks.) In like fashion, data can be stored anywhere: local device, remote device, in the cloud, distributed, etc.”) Claim 22, Kamath as modified teaches claim 1. Kamath as modified further teaches: wherein the command information is to cause the at least one end node wireless device to perform the operation comprising an operation in an automation network. (Song, para. 0073: “Referring to FIG. 1, in the AI system, at least one or more of an AI server 16, robot 11, self-driving vehicle 12, XR device 13, smartphone 14, or home appliance 15 are connected to a cloud network 10. Here, the robot 11, self-driving vehicle 12, XR device 13, smartphone 14, or home appliance 15 to which the AI technology has been applied may be referred to as an AI device (11 to 15).”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as set forth above with respect to claim 1. Claim 23, Kamath as modified teaches claim 8. Kamath as modified further teaches: wherein the command is to cause the at least one end node wireless device to turn on a light in the local environment. (Song, para. 0220: “The output unit 150 may be configured to output various types of information, such as audio, video, tactile output, and the like. The output unit 150 may include at least one of a display unit 151, an audio output unit 152, a haptic module 153, or an optical output unit 154.”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Song into Kamath, as set forth above with respect to claim 1. Claims 2 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Kamath, in view of Song, in view of Agarwal, in view of Vedal, and further in view of De la Cruz Jr, et. al. (“Jointly Pre-training with supervised, autoencoder, and value losses for deep reinforcement learning”, arXiv:1904.02206v1 [cs.LG] 3 Apr 2019; hereinafter, “de la Cruz”) Claim 2, Kamath as modified teaches claim 1. Kamath as modified further teaches:: wherein the method further comprises jointly training the autoencoder and the classifier based on a weighted sum of the first loss and the second loss. (De la Cruz sec. 3.2: “To obtain extra information in addition to the supervised features, we take inspiration from the supervised autoencoder framework which jointly trains a classifier and an autoencoder Le et al. (2018); we believe this approach will retain the important features learned through supervised pre-training and at the same time, learns additional general features from the added autoencoder loss. Finally, we blend in the value loss l o s s v s a e v with the supervised and autoencoder losses as L s a e v = L s s a e v W S F X i , y i + L a e s a e v W a e F X i , X i + L v s a e v W v F X i , x i ”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of De la Cruz into Kamath, as modified. De la Cruz teaches a pre-training strategy that jointly trains a weighted supervised classification loss, an unsupervised reconstruction loss, and an expected return loss. One of ordinary skill would have been motivated to combine the teachings of De la Cruz into Kamath, as modified, in order to enable discovery of more useful features compared to independently training in supervised or unsupervised fashion (De la Cruz, Abstract). Claim 9, Kamath as modified teaches claim 8. Kamath as modified further teaches: wherein the method further comprises jointly training the autoencoder and the classifier based on a weighted sum of the first loss and the second loss. (De la Cruz sec. 3.2: “To obtain extra information in addition to the supervised features, we take inspiration from the supervised autoencoder framework which jointly trains a classifier and an autoencoder Le et al. (2018); we believe this approach will retain the important features learned through supervised pre-training and at the same time, learns additional general features from the added autoencoder loss. Finally, we blend in the value loss l o s s v s a e v with the supervised and autoencoder losses as L s a e v = L s s a e v W S F X i , y i + L a e s a e v W a e F X i , X i + L v s a e v W v F X i , x i ”) It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of De la Cruz into Kamath, as modified, as set forth above with respect to claim 2. Claims 3 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Kamath, in view of Song, in in view of Vedal, view of García-Ordás, et. al. (“Detecting Respiratory Pathologies Using Convolutional Neural Networks and Variational Autoencoders for Unbalancing Data.” (2020) Sensors. 20. 10.3390/s20041214; hereinafter, “Garcia-Ordás”), and further in view of De la Cruz. Claim 3, Kamath as modified teaches claim 1. Kamath as modified further teaches: wherein the method further comprises: calculating the first loss according to a correlation coefficient; and (Garcia-Ordás, sec. 3.2: “For that reason, the loss function used to train a VAE is made up of two terms: ‘reconstruction term’, like in the vanilla autoencoder, that tends to make the encoder-decoder work accurately; and a ‘regularization term’ applied over the latent layer that tends to make the distributions created by the encoder close to a standard normal distribution using the Kulback-Leibler divergence” Examiner notes that Garcia-Ordás teaches a first loss using a regularization term i.e. a coefficient that correlates the distributions to normal). calculating the second loss according to a binary cross-entropy. (De la Cruz, sec. 2.5: “assume the non-optimal human actions as the true labels for each game state. The network is pre-trained with the cross-entropy loss” Examiner notes that De la Cruz teaches multi-label cross-entropy, i.e. binary cross-entropy). it would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Garcia-Ordás into Kamath, as modified. Garcia-Ordás teaches use of variational convolutional autoencoders to process sound of breaths in the form of spectrograms for diagnosing various ailments. One of ordinary skill would have been motivated to combine the teachings of Garcia-Ordás into Kamath, as modified, in order to classify the respiratory sounds into healthy, chronic, and non-chronic disease for the purpose of training a computer to classify respiratory sounds into healthy, chronic, and non-chronic disease (Garcia-Ordás, abstract). It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of De la Cruz into Kamath, as modified, as set forth above with respect to claim 2. Claim 10, Kamath as modified teaches claim 8. Kamath as modified further teaches: The at least one non-transitory computer readable storage medium of claim 1, wherein the method further comprises: calculating the first loss according to a correlation coefficient; and (Garcia-Ordás, sec. 3.2: “For that reason, the loss function used to train a VAE is made up of two terms: ‘reconstruction term’, like in the vanilla autoencoder, that tends to make the encoder-decoder work accurately; and a ‘regularization term’ applied over the latent layer that tends to make the distributions created by the encoder close to a standard normal distribution using the Kulback-Leibler divergence” Examiner notes that Garcia-Ordás teaches a first loss using a regularization term i.e. a coefficient that correlates the distributions to normal). calculating the second loss according to a binary cross-entropy. (De la Cruz, sec. 2.5: “assume the non-optimal human actions as the true labels for each game state. The network is pre-trained with the cross-entropy loss” Examiner notes that De la Cruz teaches multi-label cross-entropy, i.e. binary cross-entropy). It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of Garcia-Ordás into Kamath, as modified, as set forth above with respect to claim 3. It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the teachings of De la Cruz into Kamath, as modified, as set forth above with respect to claim 2. Responses to Applicant Remarks and Argument 35 USC §103 Starting on page 8, applicant argues that Kamath does not teach processing an audio signal that is speech information of a user. Applicant further argues that Kamath does not teach an encoder compressing real world information into a compressed spectrogram. Applicant argues that Kamath teaches an audio signal is compressed before being stored rather than compression of a spectrogram. Finally, applicant argues that one of ordinary skill in the art would not have been motivated to combine Kamath with Song. On page 9, applicant argues that Song teaches sending information from an “AI deice” to an “AI server” and that none of the references teaches joint training of an autoencoder and classifier. Finally, on page 10, applicant argues that the prior does not teach claim 8 wherein processing information identifies the presence of a person. However, after furthers search and in light of applicant’s amendments, independent claims 1, 8, and 18 are now rejected under § 103 over Kamath, Song and Agarwal. Kamath teaches and compressing both visual and audio information into a spectrogram. Upon further consideration of the limitation reciting compressing a spectrogram into a compressed spectrogram, independent claims 1, 8, and 18 are additionally rejected under Vedal. Examiner notes the compressed spectrogram is not further used or processed in independent claims therefore the compressed spectrogram was interpreted as any spectrogram which itself is a compressed representation of audio information. Finally, Agarwal further teaches such audio information being speech of a user and Kamath further teaches identification of the presence of a person. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sally T. Ley whose telephone number is (571)272-3406. The examiner can normally be reached Monday - Thursday, 10:00am - 6:00pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /STL/Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Show 8 earlier events
Nov 05, 2025
Response after Non-Final Action
Feb 23, 2026
Non-Final Rejection mailed — §103
May 15, 2026
Interview Requested
May 22, 2026
Response Filed
May 27, 2026
Applicant Interview (Telephonic)
May 27, 2026
Examiner Interview Summary
Aug 10, 2026
Final Rejection mailed — §103
Sep 24, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725021
DATA PROCESSING
4y 6m to grant Granted Sep 01, 2026
Patent 12725034
MODELLING CAUSATION IN MACHINE LEARNING
3y 11m to grant Granted Sep 01, 2026
Patent 12711349
COMPRESSION AND DECOMPRESSION OF WEIGHT VALUES
6y 4m to grant Granted Aug 18, 2026
Patent 12632746
A METHOD AND APPARATUS FOR DISPLAYING CATEGORIZED CARBON EMISSIONS
3y 6m to grant Granted May 19, 2026
Patent 12443830
COMPRESSED WEIGHT DISTRIBUTION IN NETWORKS OF NEURAL PROCESSORS
5y 9m to grant Granted Oct 14, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
23%
Grant Probability
37%
With Interview (+14.6%)
4y 9m (~2m remaining)
Median Time to Grant
High
PTA Risk
Based on 44 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month