Prosecution Insights
Last updated: October 02, 2026
Application No. 18/198,140

METHOD, SYSTEM, AND APPARATUS FOR EFFICIENT NEURAL DATA COMPRESSION FOR MACHINE TYPE COMMUNICATIONS VIA KNOWLEDGE DISTILLATION

Final Rejection §103
Filed
May 16, 2023
Priority
May 20, 2022 — provisional 63/344,418
Examiner
BREENE, PAUL J
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
9m
Est. Remaining
77%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
42 granted / 68 resolved
+6.8% vs TC avg
Strong +15% interview lift
Without
With
+15.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
14 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
27.7%
-12.3% vs TC avg
§103
47.7%
+7.7% vs TC avg
§102
8.4%
-31.6% vs TC avg
§112
15.4%
-24.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 68 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed May 18th, 2026 have been fully considered but they are not persuasive. With regard to the arguments under 35 U.S.C. § 102, the rejection is amended to ameliorate the issues raised by applicant. See Claim Rejections - 35 USC § 103 for further elaboration. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, 8-15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over US Pre Grant Patent 2021/0241108 (Chai et al; Chai) in view of “Head Network Distillation: Splitting Distilled Deep Neural Networks for Resource-Constrained Edge Computing Systems,” in IEEE Access, vol. 8, Matsubara et al; Matsubara. Regarding claims 1 and analogous claim 19: 1. A method of managing sensor data, the method comprising: receiving encoded data at a first device from a second device separate from the first device, (Chai, ¶0066) “FIG. 17 illustrates an exemplary hierarchy of computing nodes in accordance with the disclosed embodiments. This hierarchy includes a number of basic runtime engines (REs) 1701-1708, which can be located in edge devices, such as motion sensors, cameras or microphones [i.e. A method of managing sensor data, the method comprising:]. These basic REs 1701-1708 assume the existence of an associated intermediate or high-end device capable of delivering DNN models to basic REs 1701-1708 and collecting log information from basic REs 1701-1708 [i.e. receiving encoded data at a first device from a second device separate from the first device,].” 2. wherein the encoded data is generated using an artificial intelligence (Al) encoder model included in the second device based on sensor data collected by at least one sensor included in the second device; (Chai, ¶0096) “If an erroneous inference is detected (e.g. via user input or other DNN inferences), then the erroneous pathway indicates the visual features that produces the erroneous inference results. A comparison of the erroneous pathway against the activation heat map can show locations where the erroneous pathway differs from the statistical distribution of pathways in the activation heat map. To improve DNN accuracy, we can generate additional training data specifically to correct the area where there is a difference in the pathways (e.g. against the heat map) [i.e. included in the second device based on sensor data collected by at least one sensor included in the second device;]. The additional training data can be synthesized using a generative adversarial network (GAN) training methodology [i.e. wherein the encoded data is generated using an artificial intelligence (Al) encoder model].” 3. generating inference information by providing the encoded data as input to an Al inference model (Chai, ¶0097) “Hence, the above-described profiling process and the generation of the activation heat map essentially produces an explanation of how the DNN produces an inference result. The process in comparing the erroneous pathways essentially produces an explanation of how the DNN is not robust to that input data set [i.e. generating inference information by providing the encoded data as input to an Al inference model;].” 4. and performing a task based on the inference information, (Chai, ¶0139) “The context-specific models 1906 can be generated using a knowledge distillation process in a distill workflow [i.e. and performing a task based on the inference information,]. Chai does not explicitly teach 1. before the encoded data is received, wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model. Matsubara teaches: 1. before the encoded data is received, wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model. (Matsubara, pg. 7, Sect. IV, col. 1, ¶3) “We note that training is a one-time process, and the models will not be trained once they are split and deployed on the mobile device and edge server [i.e. before the encoded data is received,].” (Matsubara, pg. 7, Sect. IV.B, col. 2, ¶2) “It was empirically shown that student models trained to mimic the behavior of their teacher models (soft target) outperforms those trained on the original training dataset (hard target) in terms of prediction performance. Based on KD, in our case the output of the teacher model are used to train both the student head and tail models simultaneously, minimizing the knowledge distillation loss [i.e. wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai’s AI encoder and inference model training according to the knowledge-distillation technique taught by Matsubara, such that the encoder and inference models are jointly trained based on an output of a teacher model before deployment and receipt of encoded data. One of ordinary skill would have been motivated to make this modification to improve the efficiency of Chai’s distributed AI processing while maintaining inference performance, as Matsubara teaches that its approach improves by “limited amount of computing load to mobile devices and preserve accuracy (Matsubara, Sect. VII, col. 2, ¶2).” Regarding claim 10: 1. A device for managing sensor data, the device comprising: at least one memory storing computer-readable instructions; (Chai, ¶0045) “The computer-readable storage medium includes, but is not limited to, volatile memory, non-volatile memory, magnetic and optical storage devices such as disk drives, magnetic tape, CDs (compact discs), DVDs (digital versatile discs or digital video discs), or other media capable of storing computer-readable media now known or later developed [i.e. at least one memory storing computer-readable instructions;].” 2. and at least one processor configured to execute the computer-readable instructions to: (Chai, ¶0046) “When a computer system reads and executes the code and/or data stored on the computer-readable storage medium, the computer system performs the methods and processes embodied as data structures and code and stored within the computer-readable storage medium [i.e. and at least one processor configured to execute the computer-readable instructions to:].” 3. receive encoded data at a first device from a second device separate from the first device, (Chai, ¶0066) “FIG. 17 illustrates an exemplary hierarchy of computing nodes in accordance with the disclosed embodiments. This hierarchy includes a number of basic runtime engines (REs) 1701-1708, which can be located in edge devices, such as motion sensors, cameras or microphones. These basic REs 1701-1708 assume the existence of an associated intermediate or high-end device capable of delivering DNN models to basic REs 1701-1708 and collecting log information from basic REs 1701-1708 [i.e. receive encoded data at a first device from a second device separate from the first device,].” 4. wherein the encoded data is generated using an artificial intelligence (Al) encoder model included in the second device based on sensor data collected by at least one sensor included in the second device, (Chai, ¶0096) “If an erroneous inference is detected (e.g. via user input or other DNN inferences), then the erroneous pathway indicates the visual features that produces the erroneous inference results. A comparison of the erroneous pathway against the activation heat map can show locations where the erroneous pathway differs from the statistical distribution of pathways in the activation heat map. To improve DNN accuracy, we can generate additional training data specifically to correct the area where there is a difference in the pathways (e.g. against the heat map) [i.e. included in the second device based on sensor data collected by at least one sensor included in the second device;]. The additional training data can be synthesized using a generative adversarial network (GAN) training methodology [i.e. wherein the encoded data is generated using an artificial intelligence (Al) encoder model].” 5. generate inference information by providing the encoded data as input to an Al inference model (Chai, ¶0097) “Hence, the above-described profiling process and the generation of the activation heat map essentially produces an explanation of how the DNN produces an inference result. The process in comparing the erroneous pathways essentially produces an explanation of how the DNN is not robust to that input data set [i.e. generate inference information by providing the encoded data as input to an Al inference model;].” 6. and perform a task based on the inference information, (Chai, ¶0139) “The context-specific models 1906 can be generated using a knowledge distillation process in a distill workflow [i.e. and performing a task based on the inference information, Chai does not explicitly teach: 1. wherein before the encoded data is received wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model. Matsubara teaches: 1. wherein before the encoded data is received wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model. (Matsubara, pg. 7, Sect. IV, col. 1, ¶3) “We note that training is a one-time process, and the models will not be trained once they are split and deployed on the mobile device and edge server [i.e. before the encoded data is received,].” (Matsubara, pg. 7, Sect. IV.B, col. 2, ¶2) “It was empirically shown that student models trained to mimic the behavior of their teacher models (soft target) outperforms those trained on the original training dataset (hard target) in terms of prediction performance. Based on KD, in our case the output of the teacher model are used to train both the student head and tail models simultaneously, minimizing the knowledge distillation loss [i.e. wherein the Al encoder model and the Al inference model are jointly trained based on an output of an Al teacher model].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 2 and analogous claim 11: Chai and Matsubara teach: 1. wherein a size of the Al inference model is smaller than a size of the Al teacher model. (Chai, ¶0139) “In the distill workflow, the original model, which was developed for the cloud, serves as a teacher model, while the context-specific models 1908 are student models that learn from the teacher model.” (Chai, ¶0139) “Note that a context-specific model 1908 can be configured to have fewer parameters (e.g. less width or depth of layers) than the original model 1902, so that the context-specific model 1908 can run within constraints of the target runtime parameters [i.e. wherein a size of the Al inference model is smaller than a size of the Al teacher model].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 3 and analogous claim 12: Chai and Matsubara teach: 1. wherein a size of the encoded data is smaller than a size of the sensor data. (Chai, ¶0138) “In order to run on the smartphone, the model can go through a build process 1906 and a run process 1910 to condition it for dynamic runtime execution. The build process 1906 includes workflows for distill, compress, and compile operations, to optimize the DNN model [i.e. wherein a size of the encoded data is smaller than a size of the sensor data].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 4 and analogous claim 13: Chai and Matsubara teach: 1. obtaining a plurality of pieces of encoded data at the first device from a plurality of second devices which are separate from the first device, wherein the plurality of pieces of encoded data are generated using a plurality of Al encoder models included in the plurality of second devices; (Chai, ¶0097) “Hence, the above-described profiling process and the generation of the activation heat map essentially produces an explanation of how the DNN produces an inference result. The process in comparing the erroneous pathways essentially produces an explanation of how the DNN is not robust to that input data set. The process in producing additional data, through data collection or synthesis using GAN [i.e. obtaining a plurality of pieces of encoded data at the first device from a plurality of second devices which are separate from the first device], is essentially an adversarial training approach to make the DNN more robust based on profiling process [i.e. wherein the plurality of pieces of encoded data are generated using a plurality of Al encoder models included in the plurality of second devices;].” 2. and combining the plurality of pieces of encoded data with the encoded data to generate aggregated data, wherein the inference information is generated by the Al inference model based on the aggregated data, (Chai, ¶0097) “The process in producing additional data, through data collection or synthesis using GAN, is essentially an adversarial training approach to make the DNN more robust based on profiling process [i.e. and combining the plurality of pieces of encoded data with the encoded data to generate aggregated data, wherein the inference information is generated by the Al inference model based on the aggregated data,;].” Matsubara teaches: 1. before the plurality of pieces of encoded data are received, the plurality of Al encoder models are jointly trained with the Al encoder model and the Al inference model based on the output of the Al teacher model. (Matsubara, pg. 7, Sect. IV, col. 1, ¶3) “We note that training is a one-time process, and the models will not be trained once they are split and deployed on the mobile device and edge server [i.e. before the plurality of pieces of encoded data are received,].” (Matsubara, pg. 7, Sect. IV.B, col. 2, ¶2) “It was empirically shown that student models trained to mimic the behavior of their teacher models (soft target) outperforms those trained on the original training dataset (hard target) in terms of prediction performance. Based on KD, in our case the output of the teacher model are used to train both the student head and tail models simultaneously, minimizing the knowledge distillation loss [i.e. wherein the Al encoder model and the Al inference model are jointly trained based on the output of an Al teacher model].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 5 and analogous claim 14: Chai and Matsubara teach: 1. wherein the encoded data is quantized by the Al encoder model before being transmitted to the first device. (Chai, ¶0110) “Quantization module 112 quantizes the values for DNN parameters to reduce the memory footprint.” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 6 and analogous claim 15: Chai and Matsubara teach: 1. wherein the second device comprises a surveillance camera as the at least one sensor, and wherein the task comprises detecting at least one of an object and an event observed by the surveillance camera. (Chai, ¶0150) “FIG. 20 illustrates another example of dynamic runtime execution of the DNN model. Similar to the example in FIG. 19, in order to run on the edge devices, the model can go through a build process 2006 and a run process 2010 to condition the model for dynamic runtime execution... The context-specific model 2008 for person detection can run on a video doorbell edge device [i.e. wherein the second device comprises a surveillance camera as the at least one sensor], while the context-specific model 2008 for face recognition can run on the network edge (e.g., a content-delivery-network or CDN) [i.e. wherein the task comprises detecting at least one of an object and an event observed by the surveillance camera].” Examiner interprets the object as a person and the event as their presence in the vicinity of the video doorbell. One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 8 and analogous claim 17: Chai and Matsubara teach: 1. wherein the second device comprises an internet of things (loT) device, and wherein the encoded data is received using massive machine-type communications (mMTC). (Chai, ¶0155) “Having models move to different locations in the hierarchy of computing nodes can help track objects in motion. For example, if a tracking application has detected a blue sedan in the proximity of IOT sensors in the hierarchy of computing nodes [i.e. and wherein the encoded data is received using massive machine-type communications (mMTC).], then a specific model for blue sedans, generated in the build process 2006 from an original model 2002 [i.e. wherein the second device comprises an internet of things (loT) device,], can be deployed in the run process 2010, as described previously.” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 9 and analogous claim 18: Chai and Matsubara teach: 1. wherein the Al inference model comprises a first neural network model, and wherein the Al teacher model comprises at least one from among a second neural network model, a support vector machine (SVM) model, and an ensemble model. “The process in producing additional data, through data collection or synthesis using GAN, is essentially an adversarial training approach to make the DNN more robust based on profiling process [i.e. wherein the Al inference model comprises a first neural network model, and wherein the Al teacher model comprises at least one from among a second neural network model, a support vector machine (SVM) model, and an ensemble model].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Regarding claim 20: 1.wherein a size of the Al inference model is smaller than a size of the Al teacher model, (Chai, ¶0139) “In the distill workflow, the original model, which was developed for the cloud, serves as a teacher model, while the context-specific models 1908 are student models that learn from the teacher model [i.e. wherein a size of the Al inference model is smaller than a size of the Al teacher model,].” 2. and wherein a size of the encoded data is smaller than a size of the sensor data. (Chai, ¶0138) “In order to run on the smartphone, the model can go through a build process 1906 and a run process 1910 to condition it for dynamic runtime execution. The build process 1906 includes workflows for distill, compress, and compile operations, to optimize the DNN model [i.e. wherein a size of the encoded data is smaller than a size of the sensor data].” One of ordinary skill, at the time the invention was filed, would have been motivated to modify Chai with Matsubara. The motivation is the same as claim 1. Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over US Pre-Grant Patent 2021/0241108 (Chai et al; Chai) in view of “Head Network Distillation: Splitting Distilled Deep Neural Networks for Resource-Constrained Edge Computing Systems,” in IEEE Access, vol. 8, Matsubara et al; Matsubara in further view of US Pre-Grant Patent 2020/0357513 (Katra et al; Katra). Regarding claim 7 and analogous claim 16: Chai does not explicitly teach: 1. wherein the second device comprises a wearable device, and wherein the task comprises detecting a health event associated with a user wearing the wearable device. Katra teaches: 1. wherein the second device comprises a wearable device, and wherein the task comprises detecting a health event associated with a user wearing the wearable device. (Katra, ¶0066) “Computing device(s) 2 may interface with and/or monitor medical device(s) 6, for example, by imaging the implantation site of the medical device(s) 6, in accordance with one or more techniques of this disclosure. In addition, computing device(s) 2 may interrogate medical device(s) 6 to obtain data from medical device(s) 6, such as performance data, historical data stored to memory, battery strength of the medical device(s) 6, impedance, pulse width, pacing percentage, pulse amplitude, pacing mode, internal device temperature, etc. In some examples, computing device(s) 2 may perform an interrogation subsession with medical device(s) 6 by establishing a wireless communication with one or more of the medical device(s) 6 [i.e. and wherein the task comprises detecting a health event associated with a user wearing the wearable device.]. In some instances, medical device(s) 6 may or may not include an IMD. In an example, computing device(s) 2 interrogate the memory of a wearable medical device 6 in order to determine device operating parameters as the interrogation data [i.e. wherein the second device comprises a wearable device,].” One of ordinary skill in the art, at the time the invention was filed, would have been motivated to modify Chai and Matsubara with Katra. The motivation is to generally apply the well-known techniques of Chai and Matsubara to the UI as mentioned in Katra, as it would have been obvious to try applying a generative adversarial network to a wearable device for a patient, as “Implantation infections are estimated to occur in about 0.5% of IMD implants and about 2% of IMD replacements. Early diagnosis of IMD infections can help drive effective antibiotic therapy or device removals to treat the infection (Katra, ¶0044).” Therefore, “This non-trivial development has resulted in the UIs described herein, which are likely to provide significant cognitive and ergonomic efficiencies and advantages over previous systems going forward. The interactive and dynamic UIs include improved human-computer interactions that may provide, for a user, reduced mental workloads/burdens, improved decision-making, reduced work stress, etc (Katra, ¶0063).” Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL JUSTIN BREENE whose telephone number is (571)272-6320. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web- based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on 303-297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786 9199 (IN USA OR CANADA) or 571-272-1000. /P.J.B./Examiner, Art Unit 2129 /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

May 16, 2023
Application Filed
Feb 18, 2026
Non-Final Rejection mailed — §103
May 09, 2026
Interview Requested
May 18, 2026
Response Filed
Aug 17, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748989
DYNAMIC AUGMENTATION BASED ON DATA SAMPLE HARDNESS
6y 2m to grant Granted Sep 29, 2026
Patent 12743641
ABNORMALITY DETECTION BASED ON CAUSAL GRAPHS REPRESENTING CAUSAL RELATIONSHIPS OF ABNORMALITIES
5y 7m to grant Granted Sep 22, 2026
Patent 12737586
Methods and apparatuses for compressing parameters of neural networks
1y 4m to grant Granted Sep 15, 2026
Patent 12726184
Accelerated Learning In Neural Networks Incorporating Quantum Unitary Noise And Quantum Stochastic Rounding Using Silicon Based Quantum Dot Arrays
4y 9m to grant Granted Sep 01, 2026
Patent 12725008
TECHNIQUES FOR INPUT CLASSIFICATION AND RESPONSE USING GENERATIVE NEURAL NETWORKS
4y 5m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
77%
With Interview (+15.4%)
4y 2m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 68 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month