Prosecution Insights
Last updated: October 02, 2026
Application No. 18/674,908

COMPRESSING AND TRANSFORMING VECTOR OPERATIONS IN AN AI MODEL

Non-Final OA §101§103
Filed
May 26, 2024
Examiner
BLANCHETTE, JOSHUA B
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
48%
Grant Probability
Moderate
1-2
OA Rounds
1y 4m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 48% of resolved cases
48%
Career Allowance Rate
111 granted / 232 resolved
-12.2% vs TC avg
Strong +32% interview lift
Without
With
+31.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
37 currently pending
Career history
269
Total Applications
across all art units

Statute-Specific Performance

§101
35.2%
-4.8% vs TC avg
§103
40.2%
+0.2% vs TC avg
§102
9.9%
-30.1% vs TC avg
§112
10.9%
-29.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 232 resolved cases

Office Action

§101 §103
DETAILED ACTION Notices to Applicant This communication is a non-final rejection. Claims 1-20, as filed 05/26/2024, are currently pending and have been considered below. No priority is acknowledged. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon and the rationale supporting the rejection would be the same under either status. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 2, 6-12, 16, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang (Hongyu Wang, et al.; "BitNet: Scaling 1-bit Transformers for Large Language Models"; arXiv:2310.11453v1) in view of Hubara (Italy Hubara et al., "Binarized Neural Networks"; arXiv:1602.02505v2) and Ravi (US20210124878A1). Regarding claims 1, 11, and 20 Wang discloses: A system comprising: a processor system; and a memory that stores computer-executable instructions that are executable by the processor system (Abstract; p. 1) to at least: --convert first multi-bit components of a first vector into first single-bit components, the first vector representing a first layer in an artificial intelligence (AI) model (“we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch,” Abstract; “We first binarize the weights to either +1 or −1 with the signum function,” p. 3); --convert second multi-bit components of a second vector into second single-bit components, the second vector representing a second layer in the AI model (“As shown in Figure 2, BitNet uses the same layout as Transformers, stacking blocks of self-attention and feed-forward networks,” p. 3); … --generate a response to the AI prompt, the response comprising an output token that corresponds to a combination of a norm of the intermediate layer output multi-bit elements, a norm of the second multi-bit components, and a representation of the second layer output multi-bit elements (The layer output is the binary product rescaled by the activation norm and the weight norm; Equations 11 and 12 on p. 4; Equation 5 on p. 3; “The output activations are rescaled with {β, γ} to dequantize them to the original precision,” p. 4; “Third, we preserve the precision for the input/output embedding because the language models have to use high-precision probabilities to perform sampling,” p. 3). Wang does not expressly disclose but Hubara teaches: --convert input multi-bit components of an input vector into input single-bit components, the input vector representing an input token in an AI prompt (“We introduce a method to train Binarized Neural Networks (BNNs) – neural networks with binary weights and activations at run-time,” Abstract; “When training a BNN, we constrain both the weights and the activations to either +1 or 􀀀1. Those two values are very advantageous from a hardware perspective, as we explain in Section 4,” p. 2; This is viewed in light of Wang’s teachings: “We further quantize the activations to b-bit precision,” p. 3 and “To evaluate the capabilities with the interpretable metrics, we test both the 0-shot and 4-shot results on four downstream tasks,” p. 7); --generate first layer output multi-bit elements by combining the input single-bit components and the first single-bit components using an exclusive-or operation (“most of the 32-bit floating point multiply-accumulations are replaced by 1-bit XNOR-count operations. This could have a big impact on dedicated deep learning hardware,” p. 6; the Examiner notes that combining vectors with XNOR uses an exclusive-or operation with the output inverted.); --generate second layer output multi-bit elements by combining intermediate layer output single-bit elements, which correspond to intermediate layer output multi-bit elements that are derived from the first layer output single-bit elements, and the second single-bit components using the exclusive-or operation (Algorithm 1 on p. 3 shows that the binarized output of each layer k < L is the input to layer k + 1; “In a BNN, only the binarized values of the weights and activations are used in all calculations. As the output of one layer is the input of the next, all the layers inputs are binary, with the exception of the first layer,” P. 4). One of ordinary skill in the art would have been motivated before the effective filing date to expand Wang’s binary weight layers to include Hubara’s binary activations and XNOR counts because this would “drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations. Our estimates indicate that power efficiency can be improved by more than one order of magnitude.” Hubara p. 8. Wang does not expressly disclose but Hubara and Ravi teaches: --transform the first layer output multi-bit elements into first layer output single-bit elements by combining the first layer output multi-bit elements and a random bit sequence that is generated using a random probability distribution (Hubara binarizes each layers output before the next layer (Algorithm 1). Ravi teaches doing this with random projections, “the projection layer input 110 may be the projection network input 104 (i.e., if the projection layer 108 is the first layer in the projection network 102) or the output of another layer of the projection network 102 (e.g., a conventional layer or another projection layer),” [0057]; “When a dot product between the projection layer input and a projection vector results in a positive value, a first value may be assigned to a corresponding position in the projection function output. Conversely, when a dot product between the projection layer input and a projection vector results in a negative value, a second value may be assigned to a corresponding position in the projection function output,” [0063]; “when the random (or pseudo-random) number generators are configured to generate Normally-distributed random numbers,” [0069]). It would have been obvious to a POSITA before the effective filing date to substitute Ravi’s random-projection binarization for Hubara’s binarization step in combination with Wang at each output layer because this is a substitution of one know binarization for another with predictable results and would improve the memory savings of the combination. Additionally, it can be seen that each element is taught by either Wang, Hubara, or Ravi. The techniques of Ravi and Hubara does not affect the normal functioning of the elements of the claim which are taught by Wang. Because the elements do not affect the normal functioning of each other, the results of their combination would have been predictable. Therefore, before the effective filing date of the claimed invention, it would have been obvious to combine the teachings of Hubara and Ravi with the teachings of Wang since the result is merely a combination of old elements, and, since the elements do not affect the normal functioning of each other, the results of the combination would have been predictable. Regarding claims 2 and 12, Wang and Hubara teach: wherein the computer-executable instructions are executable by the processor system to at least: increase efficiency of the AI model by combining the input single-bit components and the first single-bit components using the exclusive-or operation to generate the first layer output multi-bit elements and combining the intermediate layer output single-bit elements and the second single-bit components using the exclusive-or operation to generate the second layer output multi-bit elements (2.3 Computational Efficiency on p. 6; this is viewed in light of Hubara’s improved power efficiency (Abstract)). The motivation to combine is the same as in claim 1. Regarding claims 6 and 16, Wang discloses: wherein the computer-executable instructions are executable by the processor system further to at least: select the output token from a plurality of tokens that are comprised in a vocabulary of the AI model as a result of the output token corresponding to the combination of the norm of the intermediate layer output multi-bit elements, the norm of the second multi-bit components, and the representation of the second layer output multi-bit elements to an extent that is greater than extents to which other tokens that are comprised in the plurality of tokens correspond to the combination of the norm of the intermediate layer output multi-bit elements, the norm of the second multi-bit components, and the representation of the second layer output multi-bit elements (“We leave the other components high-precision, e.g., 8-bit in our experiments. We summarized the reasons as follows. First, the residual connections and the layer normalization contribute negligible computation costs to large language models. Second, the computation cost of QKV transformation is much smaller than the parametric projection as the model grows larger. Third, we preserve the precision for the input/output embedding because the language models have to use high-precision probabilities to perform sampling,” p. 3). The Examiner views this teaching of Wang in light of Official Notice that selecting the vocabulary token with the highest probability (i.e., greedy decoding) was a conventional decoding choice for language models before the effective filing date. Applicant may traverse this by pointing out supposed errors, in which case, documentary evidence of this fact will be supplied. Regarding claim 7, Wang does not expressly disclose but Hubara teaches: wherein the first multi-bit components of the first vector represent first floating point numbers that are less than one; wherein the second multi-bit components of the second vector represent second floating point numbers that are less than one; and wherein the input multi-bit components of the input vector represent input floating point numbers that are less than one (algorithm 1 on p. 2; “Constrain each real-valued weight between -1 and 1, by projecting wr to -1 or 1 when the weight update brings wr outside of [-1; 1], i.e., clipping the weights during training, as per Algorithm 1. The real-valued weights would otherwise grow very large without any impact on the binary weights,” p. 4; this is viewed in like of Wang which uses “absmax quantization,” p. 3, and Wang Equation 5). The motivation to combine is the same as in claim 1. Regarding claims 8 and 17, Wang does not expressly disclose but Ravi teaches: wherein the random probability distribution is a Gaussian distribution, a Rademacher distribution, or a Bernoulli distribution (“when the random (or pseudo-random) number generators are configured to generate Normally-distributed random numbers (i.e., random numbers drawn from a Normal distribution), the values of the components of the matrices defining the projection functions are approximately Normally-distributed,” [0069]). The motivation to combine is the same as in claim 1. Regarding claim 9, Wang discloses: wherein the first multi-bit components of the first vector, the second multi-bit components of the second vector, and the input multi-bit components of the input vector are 32-bit components (“The results demonstrate the effectiveness of BitNet in achieving competitive performance levels compared to the baseline approaches, particularly for lower bit levels,” p. 9; baseline of 16 bits on p. 9; 32 bits on p. 6; this is viewed in light of Hubara’s 32-bit discussion on p. 6). The motivation to combine is the same as in claim 1. Regarding claim 10, Wang and Hubara teach using floats of different bits including 32-bit as described above with respect to claim 9. The combination does not expressly disclose 64-bit floats but this would have been an obvious design choice among a finite number of standard sizes. See MPEP 2144.05. Regarding claim 19, Wang does not expressly disclose but Hubara teaches: wherein the random probability distribution is a Bernoulli distribution (“Our second binarization function is stochastic,” p.2; algorithm reproduced below). PNG media_image1.png 60 282 media_image1.png Greyscale The motivation to combine is the same as in claim 1. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Wang (Hongyu Wang, et al.; "BitNet: Scaling 1-bit Transformers for Large Language Models"; arXiv:2310.11453v1) in view of Hubara (Italy Hubara et al., "Binarized Neural Networks"; arXiv:1602.02505v2), Ravi (US20210124878A1), and Achlioptas (Dimitris Achlioptas, "Database-friendly random projections: Johnson-Lindenstrauss with binary coins"; Journal of Computer and System Sciences 66 (2003) 671–687). Regarding claim 18, Wang does not expressly disclose but Achlioptas teaches: wherein the random probability distribution is a Rademacher distribution (Theorem 1.1 on p. 3 reproduced below). PNG media_image2.png 214 758 media_image2.png Greyscale It would have been obvious to a POSITA before the effective filing date to substitute Ravi’s random-projection binarization with Achlioptas’s Rademacher projection vectors in the combination of Wang, Hubara, and Ravi because this is a substitution of one know binarization for another with predictable results and would improve the memory savings of the combination. Additionally, it can be seen that each element is taught by either Wang, Hubara, Ravi, or Achlioptas. The techniques of Achlioptas do not affect the normal functioning of the elements of the claim which are taught by Wang, Hubara, and Ravi. Because the elements do not affect the normal functioning of each other, the results of their combination would have been predictable. Therefore, before the effective filing date of the claimed invention, it would have been obvious to combine the teachings of Hubara, Ravi, and Achlioptas with the teachings of Wang since the result is merely a combination of old elements, and, since the elements do not affect the norma Claim Objections Claims 3-5 and 13-15 are objected to for being dependent on rejected claims but would otherwise be allowable if re-written in independent form. The prior art of record does not teach or suggest, in combination with the rest of the limitations, an output token that corresponds to a combination that further includes an error estimate representing the error introduced by converting the layer vectors and the input vector into single-bit components (claims 3 and 13), nor generat[ing] the output token by multiplying the norm of the intermediate layer output multi-bit elements, the norm of the second multi-bit components, and a cosine of an angle, which takes into consideration the error estimate and a population count (or sum) of the second layer output single-bit elements (claims 4-5 and 14-15). Wang rescales the binary product by norms but has no error term and no angle. Hubara and Ravi have neither error terms nor angles. Subject Matter Eligibility Claims 1-20 are eligible under 35 U.S.C. 101. While the claims recite mathematical concepts at Step 2A Prong One (i.e., single bit conversion, exclusive-or operations, norms), at Prong Two they integrate the mathematics into a practical application by providing improvements to the functioning of the computer itself. The claims do not use a model as a generic tool to produce a result. Instead, they recite the mechanism by which the computer executes the model differently. The specification describes the improvement in a non-conclusory manner and ties it to the machine (e.g., multiple exclusive-or operations per CPU cycle). For example, [0007] states: For instance, using the single-bit elements in lieu of the multi-bit components may enable the AI model to perform inferencing using exclusive-or operations in lieu of vector multiplications, which may enable the AI model to generate the response more quickly than conventional AI models. The complexity of the computations may be reduced to an extent that enables the computations to be performed on a central processing unit, rather than a graphical processing unit. For example, multiple exclusive-or operations may be performed within a common (e.g., same) cycle of the central processing unit. Paragraph [0023] further describes that this reduction in complexity “enables the computations to be performed on a central processing unit, rather than a graphical processing unit.” Thus the claimed invention improves the functioning of the computer and amounts to a practical application at Prong Two. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Rastegari (Mohammad Rastegari er al., "XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks"; arXiv:1603.05279v4) discloses convolutional networks where layer inputs are binarized and convoluations are estimated by XNOR and bitcounting operations (Abstract). Anderson (Alexander G. Anderson et al., "The High-Dimensional Geometry of Binary Neural Networks" arXiv:1705.07199v1) teaches that binarization approximately preserves the direction of high dimensional vectors and analyzes the angle between a vector and its binarized version (Abstract). Blalock (Davis Blalock et al., "Multiplying Matrices Without Multiplying" arXiv:2106.10860v1) teaches using hashes to approximate matrix multiplication (Abstract). Jaffari (US-20190251425-A1) teaches binary neural networks in which the multiply-accumulate operation is replaced with XNOR and population count. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSHUA BLANCHETTE whose telephone number is (571)272-2299. The examiner can normally be reached on Monday - Thursday 7:30AM - 6:00PM, EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Shahid Merchant, can be reached on (571) 270-1360. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JOSHUA B BLANCHETTE/Primary Examiner, Art Unit 3624
Read full office action

Prosecution Timeline

May 26, 2024
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731674
METHOD AND CONTROL UNIT FOR CONTROLLING A MEDICAL IMAGING INSTALLATION
2y 11m to grant Granted Sep 08, 2026
Patent 12731697
Medical Procedure Preparation Guide Apparatus, Medical Procedure Preparation Guide Method, Non-Transitory Recording Medium Recording Medical Procedure Preparation Guide Program
2y 2m to grant Granted Sep 08, 2026
Patent 12718299
METHODS AND APPARATUS TO PROCESS INSURANCE CLAIMS USING CLOUD COMPUTING
3y 2m to grant Granted Aug 25, 2026
Patent 12718959
CODE FOR PATIENT CARE DEVICE CONFIGURATION
2y 6m to grant Granted Aug 25, 2026
Patent 12706186
USER INTERFACES FOR SHARED HEALTH-RELATED DATA
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
48%
Grant Probability
80%
With Interview (+31.8%)
3y 8m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 232 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month