Prosecution Insights
Last updated: October 02, 2026
Application No. 19/204,517

Federated Byte Latent Transformer for Privacy-Preserving Deep Learning

Final Rejection §101
Filed
May 10, 2025
Priority
Jun 07, 2024 — CIP of 18/737,906 +2 more
Examiner
HOANG, HAU HAI
Art Unit
2154
Tech Center
2100 — Computer Architecture & Software
Assignee
AtomBeam Technologies Inc.
OA Round
2 (Final)
78%
Grant Probability
Favorable
3-4
OA Rounds
1y 3m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
399 granted / 510 resolved
+23.2% vs TC avg
Moderate +13% lift
Without
With
+13.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
19 currently pending
Career history
537
Total Applications
across all art units

Statute-Specific Performance

§101
19.1%
-20.9% vs TC avg
§103
42.8%
+2.8% vs TC avg
§102
16.2%
-23.8% vs TC avg
§112
15.3%
-24.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 510 resolved cases

Office Action

§101
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 Claims 1, 4-12, 15-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. Claim 1 Step 1, This part of the eligibility analysis evaluates whether the claim falls within any statutory category. See MPEP 2106.03. The claim recites a system that performs at least one step. Thus, the claim is to a machine, which is one of the statutory categories of invention. (Step 1: YES). Step 2A Prong One, This part of the eligibility analysis evaluates whether the claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim "recites" a judicial exception when the judicial exception is "set forth" or "described" in the claim. Limitation: “computing for each byte in the input data a conditional entropy of a next-byte prediction distribution over a byte vocabulary conditioned on preceding bytes using a byte-level autoregressive neural network.” This limitation recites a judicial exception because it encompasses a mathematical concept. Specifically, it recites a mathematical calculation, which is defined as a mathematical operation or an act of calculating using mathematical methods to determine a variable or number. Limitation: “establishing patch boundaries when the conditional entropy of the next-byte prediction distribution exceeds a threshold.” This limitation recites a judicial exception because it encompasses a mathematical concept. Specifically, it recites a mathematical relationship or calculation involving a comparison to a threshold. “Unless it is clear that a claim recites distinct exceptions, such as a law of nature and an abstract idea, care should be taken not to parse the claim into multiple exceptions, particularly in claims involving abstract ideas.” MPEP 2106.04, subsection II.B. However, if possible, the examiner should consider the limitations together as a single abstract idea rather than as a plurality of separate abstract ideas to be analyzed individually. “For example, in a claim that includes a series of steps that recite mental steps as well as a mathematical calculation, an examiner should identify the claim as reciting both a mental process and a mathematical concept for Step 2A, Prong One to make the analysis clear on the record.” MPEP 2106.04, subsection II.B. Under such circumstances, however, the Supreme Court has treated such claims in the same manner as claims reciting a single judicial exception. Id. (discussing Bilski v. Kappos, 561 U.S. 593 (2010)). The limitations mentioned are considered together as a single abstract idea for further analysis. (Step 2A, Prong One: YES). Step 2A Prong Two: The claim recites the additional elements: a system comprising a hardware memory and one or more processors, wherein the system is configured to execute software instructions stored on nontransitory machine-readable storage media comprising software instructions that cause the system to dynamically allocate computational resources within the system based on input data by receiving input data comprising byte sequences from a plurality of client devices; segmenting the input data into patches of variable length based on the patch boundaries; encoding the patches into latent representations; transmitting the latent representations to a deep learning core; processing the latent representations using the deep learning core; and generating output data based on the processed latent representations. The limitations containing the judicial exception as well as the additional elements in the claim besides the judicial exception need to be evaluated together to determine whether the claim integrates the judicial exception into a practical application. MPEP 2106.05(a) Improvements to the Functioning of a Computer or to Any Other Technology or Technical Field: The claim does not appear to provide an improvement to the functioning of a computer or to another technology. While the claim mentions “dynamically allocate computational resources,” the claim does not provide sufficient detail to show that the claimed invention provides a technical improvement to computer operation. Instead, the claim uses the mathematical calculation of entropy to trigger standard data processing steps (segmenting, encoding, transmitting). Without a detailed description of how these steps specifically improve computer technology beyond routine processing, the claim fails this consideration. MPEP 2106.05(b) Particular Machine: The claim recites generic computer components, including a hardware memory, one or more processors, and a deep learning core. These are standard, generic computer elements used to implement the mathematical calculations and do not constitute a "particular machine" that is non-generic or specifically adapted to the judicial exception. MPEP 2106.05(c) Particular Transformation: The claim does not recite the transformation of a physical article or a specific, non-generic transformation of data that goes beyond the mere manipulation of digital information. The steps of encoding and processing data into "latent representations" are standard digital data manipulations and do not constitute a particular transformation required to integrate the exception. MPEP 2106.05(e) Other Meaningful Limitations: The additional limitations, such as “segmenting the input data into patches,” “encoding the patches,” and “transmitting the latent representations,” are generic data handling steps. They do not impose a meaningful limit on the mathematical concept of calculating entropy; rather, they are standard procedures for managing data being processed by a computer. MPEP 2106.05(g) Insignificant Extra-Solution Activity: The limitations of receiving, segmenting, encoding, transmitting, and generating output are considered insignificant extra-solution activity. These steps are standard, conventional activities required to implement a mathematical process on a computer system. They do not provide a solution to a technical problem but rather facilitate the execution of the recited abstract mathematical calculation. MPEP 2106.05(h) Field of Use and Technological Environment: The claim does not limit the application of the mathematical concept to a specific field of use or a particular technological environment that would provide a meaningful limit. The claim is directed toward general data processing using neural networks. Accordingly, the additional limitations do not integrate judicial exceptions into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B The additional elements of the claim-comprising a system of hardware memory and processors, receiving data, segmenting, encoding, transmitting, and generating output-do not, when considered individually or in an ordered combination, provide an inventive concept that adds “significantly more” than the recited mathematical concepts. Under the Berkheimer standard, these elements represent well-understood, routine, and conventional computer functions used to implement a mathematical process. Specifically, the steps of receiving data from client devices, segmenting it into patches, encoding it, and transmitting it to a core are standard, generic operations in digital data processing and machine learning environments. These steps do not represent a specific, non-generic solution to a technical problem but instead merely utilize standard computer functions to execute the abstract mathematical calculation. Therefore, the claim is patent ineligible under 35 U.S.C. § 101. Claim 4 recites “allocates more computational resources to high-entropy regions of the byte sequences and fewer computational resources to low-entropy regions.” This limitation recites an implementation of the mathematical concept of entropy into a resource management rule. The claim does not have any additional limitations that amount to significantly more than the abstract idea. Claim 5 recites “generating initial representations of elements within each patch,” “capturing contextual patterns from sequences of elements,” and “using an attention mechanism to pool element-level representations into patch-level representations.” These limitations describe the mathematical steps of an attention mechanism, involving the generation of representations. The claim does not have any additional limitations that amount to significantly more than the abstract idea. Claim 6 recites “capturing contextual patterns comprises using hash-based embeddings of element sequences of varying lengths.” Because hashing is a mathematical process used to map sequences to fixed-size values, it is a mathematical method of performing an operation on data. The claim does not have any additional limitations that amount to significantly more than the abstract idea. Claim 7 recites “the deep learning core comprises a transformer architecture that processes the latent representations without requiring fixed-vocabulary tokenization of the input data.” This limitation describes a generic machine to perform the math. The claim does not have any additional limitations that amount to significantly more than the abstract idea. Claim 10 recites “dynamically modifies patch sizes based on available computational resources while maintaining prediction accuracy.” It is a mathematical implementation of resource management and does not transform the mathematical nature of the underlying claim into a patent-eligible practical application. Claim 11 recites “is initialized using parameters from a pre-trained model and subsequently optimized for byte-level processing.” The core underlying concept for this limitation is the optimization of a mathematical model through transfer learning. The claim does not have any additional limitations that amount to significantly more than the abstract idea. Claim 12, 15-18 are similar to claims 1 and 4-7 Response to Arguments Pg. 7 Applicant argues that “…As amended, claim 1 recites "computing for each byte in the input data a conditional entropy of a next-byte prediction distribution over a byte vocabulary conditioned on preceding bytes using a byte-level autoregressive neural network" and "establishing patch boundaries when the conditional entropy of the next-byte prediction distribution exceeds a threshold," with the input data thereafter segmented "into patches of variable length based on the patch boundaries" and the resulting latent representations transmitted "to a deep learning core" for further processing. These limitations recite a specific machine having a byte-level autoregressive neural network configured to generate a next-byte prediction distribution conditioned on preceding bytes. Amended claim 1 thus is substantially more than generic computer components (a computer system, a hardware memory, and nontransitory storage media) that the examiner found insufficient under MPEP 2106.05. Further, these operations cannot characterized as steps a human could perform in the mind, because they require a trained neural network to generate a probability distribution over a byte vocabulary at each byte position before the entropy of that distribution can even be computed. See Fig. 24 & [0347] (describing the entropy model as "an innovative byte-level autoregressive language model comprising approximately 100 million parameters distributed across 14 transformer layers")…” The Applicant's argument is that the claim is patent-eligible because it uses a highly complex machine (a neural network with 100 million parameters) to perform tasks that a human could never do in their head. Also, because a human cannot calculate these complex probabilities, the steps are more than just "mental processes" or "generic computer tasks." Examiner respectfully disagrees because the core of the claim is performing a mathematical calculation (calculating “entropy”) and then comparing that result to a number (a “threshold”). These are mathematical concepts. Even if the math is incredibly difficult and requires a massive supercomputer to solve, the actual nature of the step remains a mathematical calculation. The Applicant argues that the “byte-level autoregressive neural network” makes the machine “specific.” However, simply taking an abstract idea (like a math formula) and running it on a computer does not make the idea patent-eligible. A computer, even a highly specialized one, is being used here simply as a tool to perform a math calculation. The Applicant also argues that a human cannot do this in their head; however, if a step is a mathematical calculation, it is an abstract idea. The fact that it is too complex for a human to do does not stop it from being categorized as a mathematical concept. Simply using a powerful computer to solve a math problem does not make that math problem a patentable invention. Pg. 7 - 8 Applicant argues that “… Under Step 2A, Prong Two, this ordered combination integrates the recited entropy computation into a practical application. Rather than reciting entropy calculation as a freestanding mathematical exercise, the claim uses the computed entropy for the specific, technical purpose of determining where to place patch boundaries in a byte sequence, so that computational resources can thereafter be allocated across the resulting variable-length patches according to the information density of the underlying data-smaller patches, and correspondingly more processing, for high-entropy, information-dense regions, and larger patches, with correspondingly less processing, for predictable, low-entropy regions. See Fig. 24 & [0347] ("creating smaller patches for high-entropy regions and larger patches for predictable, low- entropy regions-a capability not possible in token-based approaches with fixed vocabularies"). This is analogous to the improvement to computer functionality found sufficient in Enfish, LLC7 v. Microsoft Corp., 822 F.3d 1327 (Fed. Cir. 2016), because the claim is directed to a specific technique for improving how the underlying computer system allocates its own computational resources when processing byte-level data, rather than merely using a computer as a tool to perform an existing process…” Applicant’s argument has been considered. Examiner respectfully disagrees However, a review of Claim 1 reveals that this resource-allocation mechanism is completely missing from the claim. Claim 1 merely states that the system divides data into variable-length patches and then performs the standard steps of "transmitting" and "processing" those patches using a generic deep learning core. The claim says absolutely nothing about how the deep learning core changes its behavior, varies its processing cycles, or optimizes hardware resources based on those patches. The Applicant contends that using an entropy formula to create patches improves computer functionality. However, simply using math to create a more structured data format is a data-organization method, not a technical upgrade to the computer itself. The Applicant also argues the “capability” to process data with specialized efficiency. However, claiming a generic result-such as increased speed, reduced data size, or generalized efficiency-without explicitly claiming the unique technical software architecture that creates that result is insufficient. Because the claim alters only the mathematical logic used to segment the data stream prior to standard processing, it remains directed to an abstract mathematical concept. Pg. 8 Applicant argues that “… Even under Step 2B, the ordered combination of computing per- byte conditional entropy via a byte-level autoregressive neural network, comparing that entropy to a threshold to establish patch boundaries, segmenting accordingly, and transmitting the resulting patch representations to a deep learning core for processing without decoding is not a routine or conventional combination of generic computer functions, and amounts to significantly more than any abstract idea that might be considered recited by the claim in the aggregate…” The Applicant argues that the unconventional combination is processing the patch representations inside the deep learning core “without decoding” them first. However, claim 1 does not include the phrase “without decoding”. An applicant cannot rely on a critical feature or limitation that was left out of the literal text of the claim. The claim describes a sequence where a mathematical calculation is performed, and the resulting value is then used to segment and transmit data. Arranging these tasks in their natural, logical order does not create an inventive concept, because the underlying computer tasks (segmenting and transmitting) are being executed in a routine manner. The Applicant mentions the use of a “byte-level autoregressive neural network” and a “deep learning core” to argue that this is more than a generic computer. However, neural networks are just massive collections of mathematical equations (like the ones used to calculate the "entropy" mentioned in the claim). Using a complex math-based tool (the neural network) to perform a mathematical task (calculating entropy) does not change the fact that the core activity is a Mathematical Concept. The Applicant claims that the combination of these steps is not routine. However, the specific actions described-segmenting data into patches, transmitting data, and processing data-are fundamental, standard operations that every computer performs. Using a math result to decide when to perform these standard tasks does not make the tasks themselves significantly more than an abstract idea. Claim 12, 15-18 are similar to claims 1 and 4-7. The rejections to claims 12, 15-18 are maintained. The Applicant also mentions claim 8-9, and 19-20 reflect a specific technical means of performing computation on data while it remains encrypted, integrating those claims into a practical application under MPEP 2106.05(a) for the additional reason that they enable a technical capability (computation on encrypted data) that is not available through conventional plaintext processing. Examiner reviews the specification and takes the Applicant’s argument into his consideration Claim 1 A system comprising a hardware memory and one or more processors, wherein the system is configured to execute software instructions stored on nontransitory machine-readable storage media comprising software instructions that cause the system to dynamically allocate computational resources within the system based on input data by: receiving input data comprising byte sequences from a plurality of client devices; computing for each byte in the input data a conditional entropy of a next-byte prediction distribution over a byte vocabulary conditioned on preceding bytes using a byte-level autoregressive neural network; establishing patch boundaries when the conditional entropy of the next-byte prediction distribution exceeds a threshold; segmenting the input data into patches of variable length based on the patch boundaries; encoding the patches into latent representations by generating initial representations of elements within each patch, capturing contextual patterns from sequences of elements using hash-based embeddings of varying lengths, and using an attention mechanism to pool element-level representations into patch-level representations [0025]; encrypting the latent representations using homomorphic encryption before transmission [Claim 8]; transmitting the latent representations to a deep learning core, wherein the deep learning core comprises a latent transformer architecture configured to process latent space vectors without using an embedding layer and a positional encoding layer [0011, 0020]; processing the latent representations using the deep learning core, wherein processing comprises processing the encrypted latent representations without decoding [Claim 8]; aggregating encrypted model updates from the plurality of client devices [Claim 8]; updating the deep learning core based on the aggregated encrypted model updates [Claim 8]; implementing privacy-enhancing techniques to the encrypted model updates to prevent extraction of client device information [Claim 9], wherein implementing privacy-enhancing techniques comprises implementing differential privacy by adding calibrated noise to the encrypted model updates before aggregation, enforcing a privacy budget across multiple rounds of federated learning, and dynamically adjusting the level of noise based on the privacy budget consumption [0022]; and generating output data based on the processed latent representations. Claim 1 was rejected because it recited mathematical concepts (such as computing conditional entropy and data segmentation), which can be categorized as a judicial exception (an abstract idea). By integrating the limitations of Claims 8 and 9 along with the technical details from the specification, Claim 1 now recites a practical application that yields a concrete technological improvement to computer security and distributed resource allocation. The additions overcome the rejection for the following reasons: The math is integrated into a specific cryptographic hardware operation: The system does not merely calculate entropy in the abstract. Instead, encrypting the variable-length latent patches using homomorphic encryption prior to network transmission, and processing those patches in the core without decoding or decryption. The claim is now structurally limited to a deep learning core comprising a "latent transformer architecture" that processes latent space vectors entirely without using an embedding layer and a positional encoding layer. The system achieves this structural optimization by utilizing variable-length, hash-based embeddings coupled with an attention mechanism to pool element representations directly into patch representations. The system prevents client-side security vulnerabilities during aggregation; it aggregates encrypted model updates from distributed client nodes and implements a multi-step differential privacy protocol. It adds calibrated noise to the encrypted updates, enforces a strict privacy budget across multiple federated learning rounds, and dynamically scales the noise relative to budget consumption. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure U.S Pub 2020/0184316 – Kavukcuoglu discloses Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating discrete latent representations of input data items. One of the methods includes receiving an input data item; providing the input data item as input to an encoder neural network to obtain an encoder output for the input data item; and generating a discrete latent representation of the input data item from the encoder output, comprising: for each of the latent variables, determining, from a set of latent embedding vectors in the memory, a latent embedding vector that is nearest to the encoded vector for the latent variable. U.S. Pub 2020/0410404 - Imani discloses computing system can include a plurality of clients located outside a cloud-based computing environment, where each of the clients may be configured to encode respective original data with a respective unique secret key to generate data hypervectors that encode the original data. A collaborative machine learning system can operate in the cloud-based computing environment and can be operatively coupled to the plurality of clients, where the collaborative machine learning system can be configured to operate on the data hypervectors that encode the original data to train a machine learning model operated by the collaborative machine learning system or to generate an inference from the machine learning model. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAU HAI HOANG whose telephone number is (571)270-5894. The examiner can normally be reached 1st biwk: Mon-Thurs 7:00 AM-5:00 PM; 2nd biwk: Mon-Thurs: 7:00 am-5:00pm, Fri: 7:00 am - 4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Boris Gorney can be reached at 571-270-5626. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. HAU HAI. HOANG Primary Examiner Art Unit 2154 /HAU H HOANG/ Primary Examiner, Art Unit 2154
Read full office action

Prosecution Timeline

May 10, 2025
Application Filed
May 07, 2026
Non-Final Rejection mailed — §101
Aug 07, 2026
Response Filed
Aug 26, 2026
Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743289
LARGE OBJECT PROCESSING SYSTEM AND LARGE OBJECT PROCESSING METHOD THEREOF
3y 0m to grant Granted Sep 22, 2026
Patent 12675486
Indicator query method and system, electronic device and storage medium
1y 5m to grant Granted Jul 07, 2026
Patent 12670134
APPROXIMATE QUERY EQUIVALENCE FOR FEATURE STORES IN MACHINE LEARNING OPERATIONS PRODUCTS
2y 0m to grant Granted Jun 30, 2026
Patent 12657244
INTER-DOCUMENT ATTENTION MECHANISM
2y 1m to grant Granted Jun 16, 2026
Patent 12632429
CHARACTERIZING AND FORECASTING EVOLVING QUERY WORKLOADS
1y 5m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
78%
Grant Probability
92%
With Interview (+13.4%)
2y 7m (~1y 3m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 510 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month