Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 9 is objected to because of the following informalities: reads “processors in using a server” should read “processors using a server”. Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 9, 10-11, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over LIN (“A model-based approach to streamlining distributed training for asynchronous SGD”) in view of KAPLAN (U.S. Pub. No. US 20210304008 A1).
Regarding claim one, LIN teaches the invention substantially as claimed, including:
A method implemented by a system, wherein the method comprises: receiving, by a first node of the system, a first training subtask; ((Section 3 B, paragraph 1) We model asynchronous SGD training as the closed queueing system illustrated in Fig. 6. There are exactly K tasks, one for each worker of distributed SGD; a task models the processing of a mini-batch of examples at a worker, the trans mission of a gradient to the parameter server, its application, and the transmission of up-to-date parameters back to the worker. Each task k = 1,...,K belongs to a different class (or chain) which determines its routing among the queues of the network: task k visits the kth worker node (an infinite server, IS, station), the uplink station (a processor sharing, PS, station), the parameter server (a first-come first-served, FCFS, station), and the downlink station (another PS station), before returning to the kth worker node. Table I summarizes our notation. Let n∗ = (n∗ 1,...,n∗ K) represent the populations of task classes {1,...,K}: in our model, n∗ k = 1 for all k as each worker has exactly one task. MVA computes mean response times Tl k(n∗) for all classes k and stations l incrementally, for each n = (n1,...,nk) such that n1 = 0,...,n∗ 1, n2 = 0,...,n∗ 2, and so on, starting from Tl k(0) = 0)
While LIN does teach received a node with a subtask, it does not explicitly teach:
executing, by the first node through synchronization among first processors of the first node, the first training subtask to obtain a first weight;
However, in analogous art that similarly handles nodes and tasks, KAPLAN teaches:
executing, by the first node through synchronization among first processors of the first node, the first training subtask to obtain a first weight; ([0076] In some implementations, because each processing node of the system may continue updating the set of weights using its own local weight gradients instead of the averaged weight gradients across all processing nodes, it's possible that each processing node may end up with a different set of weight values as the training process progresses. Although the differences among the different sets of weights in the processing nodes are expected to be small (e.g., within the threshold difference), the system can improve the likelihood of convergence by performing a weight synchronization process after a predetermined number of iterations of the training process to synchronize the weight values across all the processing nodes of the neural network training system.)
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with KAPLAN’s teaching of generating a weight using a first task and, with LIN’s teaching of receiving a set of tasks and nodes, to realize, with a reasonable expectation of success, a method that nodes and correlating subtasks, as in LIN, to generate a weight, as in KAPLAN. A person of ordinary skill would have been motivated to make this combination to improve accuracy (KAPLAN [0001]).
LIN further teaches:
asynchronously receiving, from at least one second node of the system and based on a second training subtask, a second weight; ((Section 2 A, paragraph 1)To provide faster feedback and improve DNN models, clusters of distributed nodes are required. The pa rameter server [6], [23], [24] is a popular architecture to distribute the computation of SGD over multiple nodes. As depicted in Fig. 2, the training dataset is partitioned among multiple worker nodes that compute gradients in parallel, on separate mini-batches of examples (data parallelism). To synchronize their execution, worker nodes send gradients g(t) to a parameter server that holds the most up-to-date version of the weights θ. The parameter server applies the gradients and sends back the weights θ to the workers. In asynchronous SGD, weights are sent back to a worker immediately after applying its gradient;)
KAPLAN further teaches:
and obtaining, by the first node and based on the first weight and the second weight, a third weight of an artificial intelligence (Al) model. ([0045] The output data can then propagate to the next neural network layer as input to the forward propagation operation at that layer. For example, as shown in FIG. 4, forward propagation operation 402a can combine training input data with W1 weights of layer 1 to generate output data out1, which propagate to layer 2 as input. Forward propagation operation 402b can combine data out1 with W2 weights of layer 2 to generate output data out2, which can then propagate to the next layer. At the highest layer n, forward propagation operation 402n receive data outn−1 from layer n−1 (not shown in FIG. 4), combine with Wn weights of layer n, and generate output data outn.)
Regarding claim 2, KAPLAN further teaches:
The method of claim 1, further comprising executing, by the at least one second node and through synchronization among second processors of the at least one second node, the second training subtask to obtain the second weight. ([0076] In some implementations, because each processing node of the system may continue updating the set of weights using its own local weight gradients instead of the averaged weight gradients across all processing nodes, it's possible that each processing node may end up with a different set of weight values as the training process progresses. Although the differences among the different sets of weights in the processing nodes are expected to be small (e.g., within the threshold difference), the system can improve the likelihood of convergence by performing a weight synchronization process after a predetermined number of iterations of the training process to synchronize the weight values across all the processing nodes of the neural network training system.)
Regarding claim 9, KAPLAN further teaches:
The method of claim 1, further comprising synchronizing the first processors and the second processors in using a server architecture or a ring architecture.( [0149] The service provider computer may include one or more servers, perhaps arranged in a cluster, as a server farm, or as individual servers not associated with one another, and may host application and/or cloud-based software services.)
Regarding claims 11-12 and 19-20, they comprise of limitations similar to those of claims 1-2 and are therefore rejected for similar rationale. Regarding claim 18, it comprises of limitations similar to those of claim 8 and is therefore rejected for similar rationale.
Claim(s) 3-5, and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over LIN (“A model-based approach to streamlining distributed training for asynchronous SGD”), KAPLAN (U.S. Pub. No. US 20210304008 A1) in further view of WANG (U.S. Pub. No. US 20160212428 A1).
Regarding claim 3, while LIN, as modified by KAPLAN, does teach claim 2, which claim 3 is dependent upon, it does not explicitly teach:
The method of claim 2, further comprising:compressing, by the at least one second node, the second weight to obtain a compressed second weight;
However, in analogous art that similarly handles weights and nodes, WANG teaches:
The method of claim 2, further comprising:compressing, by the at least one second node, the second weight to obtain a compressed second weight; ([0024] A difference (4*dx) between the reference quantization weight B and the basic quantization weight A, a difference (4*dy) between the reference quantization weight C and the basic quantization weight A, and a difference (4*dz) between the reference quantization weight D and the basic quantization weight A can be learned in advance. The differences dx, dy and dz may be positive or negative. In the embodiment, the expanding unit 27 calculates the adjustment amounts using the differences dx, dy and dz as the basic units. And the target quantization weight to be determined can be represented as:
A+a.sub.1*dx+a.sub.2*dy+a.sub.3*dz )
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with WANG’s teaching of compressing weights and, with LIN’s, as modified by KAPLAN, teaching of receiving a set of tasks and nodes where each of the nodes are corresponding to a task and generating weights with their tasks, to realize, with a reasonable expectation of success, a method that receives nodes and correlating subtasks which generates weights, as in LIN as modified by KAPLAN, to compress a weight, as in WANG. A person of ordinary skill would have been motivated to make this combination to improve quality (WANG [0008 ]).
Lin further teaches:
and asynchronously receiving, by the first node, the compressed second weight. ; ((Section 2 A, paragraph 1)To provide faster feedback and improve DNN models, clusters of distributed nodes are required. The pa rameter server [6], [23], [24] is a popular architecture to distribute the computation of SGD over multiple nodes. As depicted in Fig. 2, the training dataset is partitioned among multiple worker nodes that compute gradients in parallel, on separate mini-batches of examples (data parallelism). To synchronize their execution, worker nodes send gradients g(t) to a parameter server that holds the most up-to-date version of the weights θ. The parameter server applies the gradients and sends back the weights θ to the workers. In asynchronous SGD, weights are sent back to a worker immediately after applying its gradient; )
Regarding claim 4, WANG further teaches:
The method of claim 3, wherein compressing the second weight comprises compressing, based a difference between a fourth weight obtained through a current synchronization among the second processors and a weight obtained through previous synchronization among the second processors, by the at least one second node, ([0024] A difference (4*dx) between the reference quantization weight B and the basic quantization weight A, a difference (4*dy) between the reference quantization weight C and the basic quantization weight A, and a difference (4*dz) between the reference quantization weight D and the basic quantization weight A can be learned in advance. The differences dx, dy and dz may be positive or negative. In the embodiment, the expanding unit 27 calculates the adjustment amounts using the differences dx, dy and dz as the basic units. And the target quantization weight to be determined can be represented as:
A+a.sub.1*dx+a.sub.2*dy+a.sub.3*dz )
KAPLAN further teaches:
the fourth weight obtained through the current synchronization among the second processors. (([0076] In some implementations, because each processing node of the system may continue updating the set of weights using its own local weight gradients instead of the averaged weight gradients across all processing nodes, it's possible that each processing node may end up with a different set of weight values as the training process progresses. Although the differences among the different sets of weights in the processing nodes are expected to be small (e.g., within the threshold difference), the system can improve the likelihood of convergence by performing a weight synchronization process after a predetermined number of iterations of the training process to synchronize the weight values across all the processing nodes of the neural network training system.))
Regarding claim 5, WANG further teaches:
The method of claim 3, wherein compressing the second weight comprises compressing, by the at least one second node based on a norm of a fourth weight of each row or each column in a fifth weight obtained through a current synchronization among the second processors, [0025] In equation (1), the weighted values a.sub.1, a.sub.2 and a.sub.3 are determined by the determination results of the approximate distance determining unit 27A. The weighted value a.sub.1 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight B, i.e., the reference quantization weight B is caused to have a greater effect on the target quantization weight. Similarly, the weighted value a.sub.2 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight C, and the weighted value a.sub.3 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight D. Based on normalization considerations, a sum of the weighted values a.sub.1, a.sub.2 and a.sub.3 may be designed as a constant value, and the output signals of the approximate distance determining unit 27A may directly be the weighted values a.sub.1, a.sub.2 and a.sub.3.
[0026] It is seen from FIG. 3(B) that, the three quantization weights located at the 1.sup.st row and the 2.sup.nd to 4.sup.th columns sequentially get closer to the reference quantization weight B, and so the respective weighted values a.sub.1 gradually increase (respectively equal to 1, 2 and 3). As the three quantization weights located at the 1.sup.st column and the 2.sup.nd to 4.sup.th rows sequentially get closer to the reference quantization weight C, the respective weighted values a.sub.2 also gradually increase (respectively equal to 1, 2 and 3). Further, taking the quantization weights located at the 4.sup.th column and the 1.sup.st to 4.sup.th rows for example, the respective weighted values a.sub.3 gets larger (respectively equal to 1, 2 and 3) as these quantization weights get closer to the reference quantization weight D. To reduce the calculation complexity level, the weighted values a.sub.1, a.sub.2 and a.sub.3 may be designed as integers.)
KAPLAN further teaches:
the fifth weight obtained through the current synchronization among the second processors. (([0076] In some implementations, because each processing node of the system may continue updating the set of weights using its own local weight gradients instead of the averaged weight gradients across all processing nodes, it's possible that each processing node may end up with a different set of weight values as the training process progresses. Although the differences among the different sets of weights in the processing nodes are expected to be small (e.g., within the threshold difference), the system can improve the likelihood of convergence by performing a weight synchronization process after a predetermined number of iterations of the training process to synchronize the weight values across all the processing nodes of the neural network training system.))
Regarding claims 12-14, they comprise of limitations similar to those of claims 3-5 and are therefore rejected for similar rationale.
Claim(s) 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over LIN (“A model-based approach to streamlining distributed training for asynchronous SGD”), KAPLAN (U.S. Pub. No. US 20210304008 A1), in further view of WANG (U.S. Pub. No. US 20160212428 A1), in further view of CAI (U.S. Pub. No. US 20180082448 A1).
Regarding claim 6, LIN further teaches:
The method of claim 1, further comprising:obtaining, by the first node, a fourth weight based on previous asynchronous update of the first node and the at least one second node; ((Section 2 A, paragraph 1) To provide faster feedback and improve DNN models, clusters of distributed nodes are required. The pa rameter server [6], [23], [24] is a popular architecture to distribute the computation of SGD over multiple nodes. As depicted in Fig. 2, the training dataset is partitioned among multiple worker nodes that compute gradients in parallel, on separate mini-batches of examples (data parallelism). To synchronize their execution, worker nodes send gradients g(t) to a parameter server that holds the most up-to-date version of the weights θ. The parameter server applies the gradients and sends back the weights θ to the workers. In asynchronous SGD, weights are sent back to a worker immediately after applying its gradient;)
While LIN, as modified by KAPLAN, does teach obtaining a weight through an asynchronous update, it does not explicitly teach:
determining, by the first node, a distance between the fourth weight and a comprehensive weight that is based on the first weight and the second weight;
However, in analogous art that similarly handle weights, WANG teaches:
determining, by the first node, a distance between the fourth weight and a comprehensive weight that is based on the first weight and the second weight; (A+a.sub.1*dx+a.sub.2*dy+a.sub.3*dz equation (1)
[0025] In equation (1), the weighted values a.sub.1, a.sub.2 and a.sub.3 are determined by the determination results of the approximate distance determining unit 27A. The weighted value a.sub.1 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight B, i.e., the reference quantization weight B is caused to have a greater effect on the target quantization weight. Similarly, the weighted value a.sub.2 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight C, and the weighted value a.sub.3 gets larger as the distance of the target quantization weight in the quantization table gets shorter to the reference quantization weight D. Based on normalization considerations, a sum of the weighted values a.sub.1, a.sub.2 and a.sub.3 may be designed as a constant value, and the output signals of the approximate distance determining unit 27A may directly be the weighted values a.sub.1, a.sub.2 and a.sub.3.)
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with WANG’s teaching of compressing weights and, with LIN’s, as modified by KAPLAN, teaching of receiving a set of tasks and nodes where each of the nodes are corresponding to a task and generating weights with their tasks, to realize, with a reasonable expectation of success, a method that receives nodes and correlating subtasks which generates weights, as in LIN as modified by KAPLAN, to compress a weight, as in WANG. A person of ordinary skill would have been motivated to make this combination to improve quality (WANG [0008 ]).
While LIN, as modified by KAPLAN and WANG, does teach getting a compressed weight of a weight achieved asynchronously, it does not explicitly teach:
and further obtaining, by the first server and based on the first weight and the second weight, the third weight when the distance is greater than a preset distance.
However, in analogous art that similarly handles server transmissions and weight calculation, CAI teaches:
and further obtaining, by the first server and based on the first weight and the second weight, the third weight when the distance is greater than a preset distance. ([0052] In some embodiments, the plurality of elements are represented by nodes in a graph, and the relations of the plurality of elements are represented by edges in the graph between the nodes. In some embodiments, the edges in the graph comprise an undirected edge indicating that elements represented by nodes connected through the undirected edge are proximity to each other.
[0053] In some embodiments, the method further comprises: in response to receiving a second input, determining a second subgraph of the graph including a second plurality of nodes, the second subgraph including the first subgraph; and causing the second subgraph to be displayed.
[0054] In some embodiments, the method further comprises: in response to receiving a third input at a third node of the second subgraph, causing an element represented by the third node to be displayed.
[0055] In some embodiments, the method further comprises: in response to receiving a fourth input at a fourth node of the second subgraph, determining a third subgraph of the graph including a third plurality of nodes in response to a distance between the fourth node and each of the third plurality of nodes being less than a predetermined distance threshold and/or a weight of each of the third plurality of nodes being greater than a predetermined weight threshold; and causing the third subgraph to be displayed.)
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with CAI’s teaching of finding the distance between weights and, with LIN’s, as modified by KAPLAN and WANG, teaching of receiving a set of tasks and nodes where each of the nodes are corresponding to a task and generating weights with their tasks and transmitting the weights asynchronously, to realize, with a reasonable expectation of success, a method that receives nodes and correlating subtasks which generates weights and transmits them asynchronously, as in LIN as modified by KAPLAN and WANG, to find a distance between weights, as in CAI. A person of ordinary skill would have been motivated to make this combination to improve system usability (CAI [0002]).
Regarding claim 15, it comprises of limitations similar to those of claim 6 and is therefore rejected for similar rationale.
Claim(s) 7-8 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over LIN (“A model-based approach to streamlining distributed training for asynchronous SGD”), KAPLAN (U.S. Pub. No. US 20210304008 A1), in further view of AKITA (U.S. Pub. No. US 20080242233 A1), in further view of HONDA (U.S. Pub. No. US 20150381913 A1).
Regarding claim 7, while LIN, as modified by KAPLAN, teaches claim 1, which claim 7 is dependent upon, it does not specifically teach:
The method of claim 1, wherein obtaining the third weight comprises:obtaining, by the first node and based on a correlation measurement function, a first correlation between the first weight and the second weight;
However, in analogous art that similarly handles weight values, AKITA teaches:
The method of claim 1, wherein obtaining the third weight comprises:obtaining, by the first node and based on a correlation measurement function, a first correlation between the first weight and the second weight; ([0091] For instance, if four weight sequence sets are assigned, the correlation between the first and the second weight sequence sets is relatively low, the correlation between the third and the fourth weight sequence sets is also relatively low )
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with AKITA’s teaching of measuring the correlation between weights and, with LIN’s, as modified by KAPLAN, teaching of receiving a set of tasks and nodes where each of the nodes are corresponding to a task and generating weights with their tasks and transmitting the weights asynchronously, to realize, with a reasonable expectation of success, a method that receives nodes and correlating subtasks which generates weights and transmits them asynchronously, as in LIN as modified by KAPLAN, to find a correlation between weights, as in AKITA. A person of ordinary skill would have been motivated to make this combination to improve performance (AKITA [0006]).
While LIN, as modified by KAPLAN and AKITA, does teach finding the correlation between weights, it does not specifically teach:
and further obtaining, by the first node and based on the first correlation, the third weight.
However, in analogous art that similarly handles weight correlation, HONDA teaches:
and further obtaining, by the first node and based on the first correlation, the third weight. ([0008] a first obtaining unit configured to obtain a first correction value using a plurality of the selected pixel values; a second obtaining unit configured to obtain a weight for the first correction value based on a magnitude of a first correlation value)
It would have been obvious to a person skilled in the art before the effective filing date of the invention to have combined with HONDA’s teaching of finding a weight using correlation and, with LIN’s, as modified by KAPLAN and AKITA, teaching of receiving a set of tasks and nodes where each of the nodes are corresponding to a task and generating weights with their tasks and transmitting the weights asynchronously, to realize, with a reasonable expectation of success, a method that receives nodes and correlating subtasks which generates weights and transmits them asynchronously, as in LIN as modified by KAPLAN and AKITA, to find a weight using a correlation, as in HONDA. A person of ordinary skill would have been motivated to make this combination to improve system accuracy (HONDA [0005]).
Regarding claim 8, AKITA further teaches:
The method of claim 7, further comprising:obtaining, by the first server based on the correlation measurement function, a second correlation between a fourth weight and a comprehensive weight that is based on the first weight and the second weight, (([0091] For instance, if four weight sequence sets are assigned, the correlation between the first and the second weight sequence sets is relatively low, the correlation between the third and the fourth weight sequence sets is also relatively low ))
LIN further teaches:
wherein the fourth weight is based on previous asynchronous update of the first node and the at least one second server; ((Section 2 A, paragraph 1)To provide faster feedback and improve DNN models, clusters of distributed nodes are required. The pa rameter server [6], [23], [24] is a popular architecture to distribute the computation of SGD over multiple nodes. As depicted in Fig. 2, the training dataset is partitioned among multiple worker nodes that compute gradients in parallel, on separate mini-batches of examples (data parallelism). To synchronize their execution, worker nodes send gradients g(t) to a parameter server that holds the most up-to-date version of the weights θ. The parameter server applies the gradients and sends back the weights θ to the workers. In asynchronous SGD, weights are sent back to a worker immediately after applying its gradient;)
KAPLAN further teaches:
obtaining, by the first server and based on the second correlation, a difference between the comprehensive weight and the fourth weight, and a first variation of previous global update, a second variation of current global update; ([0069] Once processing node 704-1 has obtained the next set of weights (e.g., either by receiving the next set of weights in a communication from another processing node, or by calculating the next set of weights itself), processing node 704-1 can compare the set of speculative weights generated from the local weight gradients to the set of weights generated from the averaged weight gradients to determine the difference between the two sets of weights. In some implementations, the difference of the two sets of weights can be determined, for example, as the average of the differences between corresponding weight values from the two sets, the median of the differences between corresponding weight values from the two sets, the maximum of the differences between corresponding weight values from the two sets, or the root mean square of the differences between corresponding weight values from the two sets, etc. In other implementations, other metrics can be used to measure the difference between the two sets of weights.
[0070] The gradients exchange and weights comparison process 716-1 may then compare the difference between the two sets of weights with a threshold difference to determine how close the set of speculative weights is to the actual weights calculated from the averaged weight gradients. If the difference between the two sets of weights is at or below the threshold difference, then the speculative weights are close enough to the actual weights, and the training process can continue using the results obtained from performing the second iteration of the training process with the set of speculative weights. In other words, processing node 704-1 can continue training the neural network model with the results obtained from using the set of speculative weights in forward propagate 722-1 and backward propagation 724-1, and proceed with further iterations of the training process.
[0071] As shown in FIG. 7, the training process may continue by performing a local weights update 728-1 using the set of local weight gradients determined form backward propagation 724-1 to enable initiation of the third iteration of the training process on the batch 3-1 input data set while waiting for the next set of weights to become available from the gradients exchange and weights comparison process 726-1. The training process may repeat for any number of iterations until the error of the neural network model is lowered to a certain error threshold across the training data set, or reaches a minimum error.) and further obtaining, by the first node and based on the second variation, the third weight. ([0096] When weights update module 1116 receives the local weight gradients from neural network computation circuitry 1112, weights update module 1116 may compute the speculative weights using the local weight gradients and send the speculative weights as updated weights to neural network computation circuitry 1112 such that training on the next batch of data can be started. When processor 1114 receives a global weights update, weights update module 1116 can compare the difference between the weight values from the global weights update with the speculative weight values. If the difference is at or below a threshold difference, the training process can continue using the speculative weights. If the difference exceeds a threshold difference, weights update module 1116 may send the weight values from the global weights update weights to neural network computation circuitry 1112 and restart the training iteration using the weights from the global weights update.)
Regarding claims 16-17, they comprise of limitations similar to those of claim 7-8 and are therefore rejected for similar rationale.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SKIELER A KOWALIK whose telephone number is (571)272-1850. The examiner can normally be reached 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D Reyes can be reached at (571)270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SKIELER ALEXANDER KOWALIK/Examiner, Art Unit 2142
/Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142