DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/13/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 4-5, 8-9, 11-12, 15-16 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. ("Online trajectory recovery from offline handwritten
Japanese kanji characters of multiple strokes") in view of Solomon et al. US PG-Pub(US 20220188542 A1).
Regarding Claim 1, Nguyen teaches a computer-implemented method(Abstract—We propose a deep neural network-based method to recover dynamic online trajectories from offline handwritten Japanese kanji character images) comprising: generating, from an image of hand-written text, trajectory data for the hand-written text using a first machine learning model that provides, as output, the trajectory data (Page 8323, IV. Training Stage, “This section expresses the details of data preparation and configurations to train our proposed network. Since there is no Japanese handwriting database collecting both online and offline patterns at the same time, we generate offline handwritten character images from online handwritten character patterns. Although the rendered images are not real offline patterns, they are useful”, this section of the prior art discloses training a encoder/decoder network to process the handwritten text and obtain the trajectory of the handwritten text.); provides, as output, machine-readable encoding that represents the hand-written text(Page 8322, Left Col, Paragraph 1, “Fig. 2 shows our proposed encoder decoder model for online trajectory recovery from handwritten character images.” Figure 2 discloses a encoder/decoder model used to generate a machine readable encoding of the hand-written text output. ).
Nguyen does not explicitly teach aligning, using a software module, the trajectory data with image data of the image to generate an aligned data set; and inputting the aligned data set into a second machine learning model.
Solomon teaches aligning, using a software module, the trajectory data with image data of the image ([0038] An image-forming component 120 renders the spatial cluster into an image, referred to as image information herein. That is, the image information represents the ink strokes in a spatial cluster using pixels, whereas the spatial information represents the ink strokes as a series of positions traversed by the user in drawing the ink strokes. [0040] A spatial data encoder 122 maps the spatial information to a first feature embedding 124 within a distributed feature space, while an image data encoder 126 maps the image information into a second feature embedding 128 within the same feature space. In one non-limiting implementation, the spatial data encoder 122 can be implemented by a first convolutional neural network, while the image data encoder 126 is implemented by a second convolutional neural network., ¶[0038] discloses spatial information of the handwritten text is obtained and ¶[0040] discloses mapping the trajectory/handwritten text into a first embedding.); to generate an aligned data set (¶[0043], “A fusion component 130 processes the first feature embedding 124 and the second feature embedding 128, to generate an output feature embedding 132. ” discloses combining/aligning the feature embedding to generate a final output that is aligned.)and inputting the aligned data set into a second machine learning model.(¶[0044] discloses taking the output feature embedding and performing a classification using a second classifier.)
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the claimed invention as taught by Nguyen with Solomon in order to align the trajectory data and handwritten data. One skilled in the art would have been motivated to modify Nguyen in this manner in order to identify a grouping of ink strokes created by the user. (Solomon, Abstract)
Regarding Claim 4, the combination of Nguyen and Solomon teach the computer-implemented method of claim 1, where Nguyen further teaches wherein the image data comprises multiple pixels of the image and wherein the aligning the trajectory data with the image data comprises correlating each pixel of the multiple pixels with a corresponding portion of the trajectory data. (Page 8324, Evaluating Stage, A. Evaluation methods for online trajectory recovery, Paragraph 2, “In this paper, we present two evaluation approaches for our proposed trajectory recovery method. The first approach is visual verification, and the second one is recognition performance evaluation by handwritten character recognition. In the first approach, we show some examples of both successful and unsuccessful samples, which are useful to consider how to revise the method in further research. The second approach points out a quantitative measurement to determine whether our proposed method can be used to improve offline handwritten character recognition.”, discloses aligning the trajectory data with image data acquired as shown in figure 3 shows the recovered handwritten text.)
Regarding Claim 5, the combination of Nguyen and Solomon teach the computer-implemented method of claim 4, where Nguyen further teaches wherein the aligning comprises performing the correlating to generate a combined grid and inputting the combined grid into a convolutional filter that, in response, produces a feature map that is the aligned data set that is input into the second machine learning model. (Page 8325, C. Evaluation by handwritten character recognizer, Paragraph 1,” To evaluate the quality of reconstructed trajectories, we employed offline and online recognizers. The offline handwritten character recognizer is composed of 4 Convolutional layers with kernel sizes of 3x3 and depths of 100, 200, 300, 400, respectively. Each Convolutional layer is followed by a Max Pooling layer. After the fourth Max Pooling layer, a Fully Connected layer of 500 ReLU cells with a dropout rate of 0.25 is employed. Finally, a classification layer based on Softmax activation function is used to compute the recognition probabilities of a converted offline pattern.” Discloses evaluating the quality of the reconstructed trajectories and inputting the trajectory into a recognized with a convolutional filter that outputs a feature map.)
Regarding Claim 8, the combination of Nguyen and Solomon teach the computer-implemented method of claim 1, where Nguyen further teaches further comprising training the first machine learning using a training data set that correlates training image data with training trajectory data. (Page 8323, IV. Training Stage, “This section expresses the details of data preparation and configurations to train our proposed network. Since there is no Japanese handwriting database collecting both online and offline patterns at the same time, we generate offline handwritten character images from online handwritten character patterns. Although the rendered images are not real offline patterns, they are useful”, this section of the prior art discloses training a encoder/decoder network to process the handwritten text and obtain the trajectory of the handwritten text.);
Regarding Claim 9, the combination of Nguyen and Solomon teach the computer-implemented method of claim 1, where Nguyen further teaches wherein the first machine learning model comprises a sequence prediction model (Page 8322, LSTM-Based decoder, First Paragraph “The LSTM-based decoder is the state-of-the-art method for generating a sequence of predictions from encoded features”, LSTM-based decoder is a sequence prediction model.).
Regarding Claim 11, the combination of Nguyen and Solomon teach the computer-implemented method of claim 1, where Nguyen further teaches wherein the trajectory data comprises stroke end flags corresponding to respective points of the hand-written text(Page 8323, IV. TRAINING STAGE, A. Data preparation, Paragraph 1 discloses determining stroke order of the handwritten text.), the stroke end flags respectively indicating whether a writing utensil used to write the hand-written text was linked to an immediately subsequent point of the hand-written text or was lifted up at the respective point. (Page 8326, Left Col, Last Paragraph, “line recognizer for their offline handwritten character images but correctly recognized by the combined recognizer for their offline and the reconstructed online character patterns. In Fig 8 (a) and (b), the kanji characters under the offline images colored in black are the recognition results by the offline recognizer, while those under the reconstructed online trajectories colored in red are recognition results by the combined recognizer. In the case in Fig. 8 (a), two kanji characters recognized by the off-line recognizer and the combined recognizer look similar, especially when we only consider the offline image. However, their online trajectories are different, especially the number of strokes. Thus, the reconstructed online trajectory helps the combined recognizer not misrecognize it. In Fig. 8 (b), some strokes are written in touching, which causes misrecognition when using only the offline image. Due to the correctly reconstructed online trajectory with their separated strokes, the combined recognizer correctly recognizes it.”, this section discloses using a recognizer to determine the number of strokes in the handwritten text to measure trajectory.)
Regarding Claim 12, claim 12 is considered a system claim substantially corresponding to claim 1. Please see the discussion of claim 1 above for a discussion of similar limitations. Furthermore, Nguyen teaches a computer system(Fig. 2. Proposed encoder-decoder network with an attention layer to reconstruct an online trajectory from an offline image.) a processor set(Page 8323, B. Configurations for training network, Paragraph 1 discloses a GPU to perform the tasks of image processing.); one or more computer-readable storage media; and program instructions stored on the one or more storage media to cause the processor set to perform operations comprising (See, Page 8323, B. Configurations for training network, Paragraph 1 discloses a GPU which would inherently be coupled to a memory to perform the tasks of image processing.)
Regarding claim 15, it is substantially similar to claim 4 respectively, and is rejected in the same manner, the same art, and reasoning applying.
Regarding claim 16, it is substantially similar to claim 5 respectively, and is rejected in the same manner, the same art, and reasoning applying.
Regarding claim 19, it is substantially similar to claim 8 respectively, and is rejected in the same manner, the same art, and reasoning applying.
Regarding Claim 20, claim 20 is considered a computer program product claim substantially corresponding to claim 1. Please see the discussion of claim 1 above for a discussion of similar limitations. Furthermore, Nguyen teaches a computer program product comprising: one or more computer-readable storage media; and program instructions stored on the one or more storage media to perform operations comprising: (See, Page 8323, B. Configurations for training network, Paragraph 1 discloses a GPU which would inherently be coupled to a memory to perform the tasks of image processing.)
Claims 2-3 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. ("Online trajectory recovery from offline handwritten
Japanese kanji characters of multiple strokes") in view of Solomon et al. US PG-Pub(US 20220188542 A1) in view of Huang et al. ("Fully Decoupling Trajectory and Scene Encoding for Lightweight Heatmap-Oriented Trajectory Prediction").
Regarding Claim 2, while the combination of Nguyen and Solomon teach the computer-implemented method of claim 1, they do not explicitly teach wherein the aligning the trajectory data with the image data comprises aligning multiple image embeddings of the image data with multiple trajectory embeddings of the trajectory data using cross attention in a transformer machine learning model of the software module.
Huang teaches wherein the aligning the trajectory data with the image data comprises aligning multiple image embeddings of the image data with multiple trajectory embeddings of the trajectory data using cross attention in a transformer machine learning model of the software module. (Page 9144, Left Col, Paragraph 1, “where (1) trajectory self attention further extracts the temporal information in trajectory features, (2) trajectory-to-image cross attention aims to query useful scene information from image features and update trajectory features and finally, (3) image-to-trajectory cross attention fuses useful trajectory features into image features. After stacking these attentions several times, the updated image features would contain rich trajectory information and we send them into a scale-up convolution to generate high-resolution heatmaps for endpoints”, discloses using cross attention to align the image and trajectory embeddings)
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the claimed invention as taught by Nguyen and Solomon with Huang in order to align the image embedding the trajectory embedding using cross attention. One skilled in the art would have been motivated to modify Nguyen and Solomon in this manner in order to propose a transformer-based heatmap decoder to model the complex interaction between high-level trajectory and image features via trajectory self-attention. (Huang, Abstract)
Regarding Claim 3, the combination of Nguyen, Solomon and Huang teach the computer-implemented method of claim 2, Nguyen further teaches further comprising training the transformer machine learning model using contrastive loss, wherein the transformer machine learning model that performs the aligning comprises the trained transformer machine learning model. (Page 8323, Left Col, Last paragraph, “In order to compute the reconstruction loss from GMM parameters, we use Eq. (7) and Eq. (8): where Np is the number of online points of the target trajectory. Note that the loss Lp is computed at all timesteps, while Ls is only computed depending on the length of the online trajectory (Np). The general loss L is the sum of Ls and Lp. It is used to optimize the encoder-decoder network”, discloses training the machine learning model by calculating the reconstruction loss.)
Regarding claim 13, it is substantially similar to claim 2 respectively, and is rejected in the same manner, the same art, and reasoning applying.
Regarding claim 14, it is substantially similar to claim 3 respectively, and is rejected in the same manner, the same art, and reasoning applying.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. ("Online trajectory recovery from offline handwritten
Japanese kanji characters of multiple strokes") in view of Solomon et al. US PG-Pub(US 20220188542 A1) in view of Bhunia et al. ("Handwriting Trajectory Recovery using End-to-End Deep Encoder-Decoder Network").
Regarding Claim 10, while the combination of Nguyen and Solomon teach the computer-implemented method of claim 9, they do not explicitly teach wherein the generating the trajectory data comprises: generating an encoded representation of the image; and initializing a hidden state of the sequence prediction model using the encoded representation of the image.
Bhunia teaches wherein the generating the trajectory data comprises: generating an encoded representation of the image; and initializing a hidden state of the sequence prediction model using the encoded representation of the image. (Abstract, “The proposed encoder module consists of Convolutional LSTM network, which takes an offline character image as the input and encodes the feature sequence to a hidden representation. The output of the encoder is fed to a decoder LSTM and we get the successive coordinate points from every time step of the decoder LSTM”, the character image is encoded into a hidden state which is fed into a decoder.)
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify the claimed invention as taught by Nguyen and Solomon with Bhunia in order to generate an encoded image and initialize a hidden state of the sequence model. One skilled in the art would have been motivated to modify Nguyen and Solomon in this manner in order to introduce a novel technique to recover the pen trajectory of offline characters which is a crucial step for handwritten character recognition. (Bhunia, Abstract)
Allowable Subject Matter
Claims 6,7 and 17-18 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claims 6 and 17, the primary reason for the allowance of the claims is the inclusion of the limitations, “wherein the aligning comprises: inputting the trajectory data into a first grid; inputting the image data into a color grid; inputting the first grid and the color grid separately into one or more convolutional filters so that, in response the one or more convolutional filters output a first feature map and a second feature map, respectively, and combining the first feature map and the second feature map to produce the aligned data set that is input into the second machine learning model.”, in all the claims which is not found in the prior art references. It is noted that the examiner has not found any other prior art to anticipate or obviate the quoted claim limitations supra, when read in light/combination of the other claimed limitations within the cited claims. Also, it is noted that the quoted limitations, in combination with the other claim limitations of the cited claims, deem the claims patentable, not just the consideration of the quoted limitations by themselves.
Claims 7 and 18 are allowable by virtue of dependency on claim 6 and 17.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HAN D HOANG whose telephone number is (571)272-4344. The examiner can normally be reached Monday-Friday 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JOHN M VILLECCO can be reached at 571-272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAN HOANG/Primary Examiner, Art Unit 2661