Prosecution Insights
Last updated: October 02, 2026
Application No. 18/199,611

SYSTEMS AND METHODS FOR SELF-SUPPERVISED TIME-SERIES REPRESENTATION LEARNING

Final Rejection §101§103§112
Filed
May 19, 2023
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Royal Bank of Canada
OA Round
2 (Final)
51%
Grant Probability
Moderate
3-4
OA Rounds
1y 0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
41 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on June 10, 2026, in which claims 1, 3, 9, and 11 are currently amended. Claims 5 and 8 are canceled. Claims 1-4, 6, 7, and 9-13 are currently pending. Specification Examiner’s objection to the specification is maintained in view of the misspelling “suppervised” which should be “supervised”. Response to Arguments The previous rejections to claims 1-13 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-13 under 35 U.S.C. 101 based on amendment have been considered. Applicant’s arguments directed towards claims 1, 3-4, 6, 7, 9-11, and 13 are persuasive. The rejections to claims 1, 3-4, 6, 7, 9-11, and 13 under 35 U.S.C. § 101 are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Claim 12 is directed towards a neural network which is interpreted as software-per-se under broadest reasonable interpretation and the rejection should be maintained. Applicant’s arguments with respect to rejection of claims 1-13 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Specification The following title is suggested: “Systems and methods for self-supervised time-series representation learning”. Claim Rejections - 35 USC § 101 101 Rejection 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 12 is rejected under 35 USC § 101 because the claimed invention is directed to non-statutory subject matter. Regarding claims 12, claims 12 is directed towards non-statutory subject matter, “software per se”. Claim 12 is directed towards a neural network, which in view of the instant specification and what is known in the art is interpreted as software ([¶0028] “The CPU 104 performs arithmetic calculations and control functions to execute software stored in a non-transitory internal memory”). Therefore, claim 12 is rejected as software-per-se. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 6 and 7 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claims 6 and 7, claims 6 and 7 are dependent on cancelled claim 5. For this reason the scope of claims 6 and 7 cannot be determined and therefore, are indefinite. In the interest of further examination claims 6 and 7 are interpreted as being dependent on claim 1. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3, 4, 6, 9, 11, 12, and 13 are rejected under U.S.C. §103 as being unpatentable over the combination of Woo (US20220261651A1) and Zheng (“ReSSL: Relational Self-Supervised Learning with Weak Augmentation”, 2021). Regarding claim 1, Woo teaches A method of generating a trained neural network that provides a universal time-series representation, the method comprising:([¶0003] "to use ample amounts of unlabeled data to learn an effective and general representation of the time series data to perform tasks like classification, even when only limited amounts of labeled data are available." [¶0008] "FIG. 5 is a table showing statistics of time series classification datasets. Train, validation, and test splits represent the number of samples in each set after performing a 50/25/25 split on all available samples.") applying a first augmented subsequence of an input time-series ([¶0065] "Random augmentations applied in the implementations include […] window slicing. [...] In window Slicing, given a window size w and time series (x1, . . . xdw), where d is a hyperparameter, a subsequence of length w from the time series is randomly sampled") to a teacher encoder ([¶0031] "the augmented versions 304 of two time series data from samples 302 are fed to an encoder 306 to generate vector representations 308 including queries and keys" [¶0028] "multi-view contrastive relational learning framework 130 may consider L views [...] Multi-view contrastive relational learning framework 130 may include an encoder for each view [...] Each encoder may encode time series data for a view and generate a representation of the view [...] the first view, V1 may be the main view (identity transformation)" [¶0037] " the multi-view information is transferred between the main view 310 and auxiliary views 312 via inter-sample structure") to generate a teacher representation of the input time-series;([¶0051] "a main encoder may generate a first representation of a first transformed data sample of a training data sample from the training dataset" Main encoder corresponding to main view (view 1) interpreted as teacher encoder to generate a teacher representation of the input time-series.) applying a second augmented subsequence of the input time-series to a student encoder to generate a student representation of the input time-series; ([¶0033] "two data augmentation operators are sampled from T (t˜T and t′˜T)" [¶0051] "At process 430, one or more auxiliary encoders may generate, in parallel to the main encoder, one or more auxiliary representations of different transformed version of the training data sample" Auxiliary encoder interpreted as a student encoder to generate a student representation of the input time-series) determining teacher instance similarities by comparing a representation of the teacher representation at a first temporal location to representations of a plurality of anchor sequences at the first temporal location;([¶0038] "P(kj l|qi l) is the probability of j-th key being matched to the i-th query, both of view l" [¶0065] "In window Slicing, given a window size w and time series (x1, . . . xdw), where d is a hyperparameter, a subsequence of length w from the time series is randomly sampled. In a particular example, d=2 for experiments using window slicing is selected" Woo calculates a teacher/main-view similarity between qiI and each of a plurality of keys. Woo explicitly anticipates the augmented samples being part of a sliced window / temporal location) determining a temporal loss based on teacher temporal similarities between representations at different temporal locations within the teacher representation of the input time-series and student temporal similarities between representations at different temporal locations within the student representation of the input time-series; ([¶0027] "training set X may be segmented into nonoverlapping windows of length W" [¶0024] "the single view contrastive learning module is configured to compute a contrastive loss component based on similarities among one or more individual views, such as one or more auxiliary representations" Woo explicitly maps a single long time series into separate windows, with window I corresponding to one temporal portion and window j corresponding to another. Its mathematical definition makes this particularly clear: window I contains samples starting at (t+iW) and ending around (t+(i+1)W). Woo then performs contrastive learning separately at the respective view encoders and computes losses based on similarities between sample representations.) determining an instance loss based on teacher instance similarities between representations at common temporal locations within the teacher representation and a plurality of anchor representations and student instance similarities between representations at common temporal locations within the student representation and the plurality of anchor representations; ([¶0037] "This transfer of the inter-sample relationships between the auxiliary views 312 and the main view 310 is realized by using the normalized similarities 320 between queries and keys 318 as distributions in a cross-entropy loss 322." Eqn. 2 of Woo explicitly compares the two view-dependent distributions over the same sample index i and plurality j=1 to K.) updating the student encoder based on the temporal loss and instance loss; and([¶0052] "the one or more auxiliary encoders may be updated based on a combination of contrastive loss component and the relational loss component or a combined loss objective via backpropagation. The combined loss objective may be computed by combining the contrastive loss component and the relational loss component."). However, Woo does not explicitly teach updating the teacher encoder as a moving average of the student encoder.. Zheng, in the same field of endeavor, teaches updating the teacher encoder as a moving average of the student encoder.([p. 4] "the teacher is updated with a "momentum update" (exponential moving average) of the student"). Woo as well as Zheng are directed towards self supervised relational representation learning. Therefore, Woo as well as Zheng are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Woo with the teachings of Zheng by updating the teacher encoder as a moving average of the student encoder. Zheng provides as additional motivation for combination ([p. 4] "the quality of the target similarity distribution p is crucial, to make the similarity distribution reliable and stable, we usually require a large batch size which is very unfriendly to GPU memories. To resolve this issue, we utilize a "momentum update" network"). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 3, the combination of Woo and Zheng teaches The method of claim 1, wherein the first augmented subsequence is generated by applying a first augmentation to a first sampled subsequence of the input time series and the second augmented subsequence is generated by applying a second augmentation to a second sampled subsequence of the input time series, (Woo [¶0033] "two data augmentation operators are sampled from T (t˜T and t′˜T)" [¶0051] "At process 430, one or more auxiliary encoders may generate, in parallel to the main encoder, one or more auxiliary representations of different transformed version of the training data sample" Woo explicitly selects two augmentation operators t and t' for the same time-series training sample and separately transforms (x_i_ with each operator. Woo's views are likewise defined as a "specific transformation" of the input time-series data.) wherein the first and second sampled subsequences have a minimum overlap.(Woo [¶0051] "a main encoder may generate a first representation of a first transformed data sample of a training data sample from the training dataset. For example, the first transformed data sample is identity transformation of the training data sample. At process 430, one or more auxiliary encoders may generate, in parallel to the main encoder, one or more auxiliary representations of different transformed version of the training data sample, respectively" Woo's first and second underlying sampled subsequences are the same window, their overlap is therefore 100% which necessarily exceeds any minimum overlap requirement.). Regarding claim 4, the combination of Woo and Zheng teaches The method of claim 3, wherein the first augmentation and the second augmentation have the same number of timestamps.(Woo [¶0051] "a main encoder may generate a first representation of a first transformed data sample of a training data sample from the training dataset. For example, the first transformed data sample is identity transformation of the training data sample. At process 430, one or more auxiliary encoders may generate, in parallel to the main encoder, one or more auxiliary representations of different transformed version of the training data sample, respectively" [¶0057] "Two long time series datasets, MIT-BIH Atrial Fibrillation (AFib) and Smartphone Based Recognition of Human Activities and Postural Transitions Data Set (HAPT) were further evaluated. AFib contains 25 ECG recordings, each 10 hours in duration, with two ECG signals sampled at 250 samples per seconds. Each time stamp is annotated with one of four labels. Windows of length 2500 were taken for AFib. HAPT was collected from subjects wearing waist-mounted smartphones performing daily activities, with a total of 12 class labels. The dataset consists of 6 readings per time stamp, representing the raw 3-axis accelerometer and gyroscope readings. Windows of length 50 were taken for HAPT."). Regarding claim 6, the combination of Woo and Zheng teaches The method of claim 5, further comprising determining the student temporal similarities by: comparing a representation of the student representation at the particular temporal location to representations of the student representation at other temporal locations.(Woo [¶0033] "Given a sample, a positive pair is an augmented version of that sample, while any other sample from the dataset may be used as a negative sample. Specifically, the instance discrimination task may define a family of data augmentations T. For a single time series data xi, two data augmentation operators are sampled from T (t˜T and t′˜T)" [¶0035] "there are K negative embeddings" [¶0052] "a contrastive loss component may be computed based on similarities among the one or more auxiliary representations that are generated from a same encoder"). Regarding claim 9, the combination of Woo and Zheng teaches The method of claim 1, further comprising determining the student instance similarities by: comparing a representation of the student representation at a second temporal location to representations of the plurality of anchor sequences at the second temporal location.(Woo [¶0033] "Given a sample, a positive pair is an augmented version of that sample, while any other sample from the dataset may be used as a negative sample. Specifically, the instance discrimination task may define a family of data augmentations T. For a single time series data xi, two data augmentation operators are sampled from T (t˜T and t′˜T)" [¶0035] "there are K negative embeddings" [¶0052] "a contrastive loss component may be computed based on similarities among the one or more auxiliary representations that are generated from a same encoder"). Regarding claim 11, the combination of Woo and Zheng teaches The method of claim 10, wherein the plurality of anchor sequences comprise subsequences of previously processed teacher representations(Woo [¶0027] "multi-view contrastive relational learning framework 130 may be trained to learn a representation for a sub-sequence or a window of single long time series data" [¶0031] "the augmented versions 304 of two time series data from samples 302 are fed to an encoder 306 to generate vector representations 308 including queries and keys" [¶0028] "multi-view contrastive relational learning framework 130 may consider L views [...] Multi-view contrastive relational learning framework 130 may include an encoder for each view [...] Each encoder may encode time series data for a view and generate a representation of the view [...] the first view, V1 may be the main view (identity transformation)"). Regarding claim 12, the combination of Woo and Zheng teaches A neural network trained according to the method of claim 1.(Woo [¶0003] "to use ample amounts of unlabeled data to learn an effective and general representation of the time series data to perform tasks like classification, even when only limited amounts of labeled data are available." [¶0008] "FIG. 5 is a table showing statistics of time series classification datasets. Train, validation, and test splits represent the number of samples in each set after performing a 50/25/25 split on all available samples."). Regarding claim 13, the combination of Woo and Zheng teaches A non-transitory computer readable memory storing instructions, which when executed by a processor of a system configure the system to perform the method of claim 1.(Woo [¶0023] "memory 120 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 110) may cause the one or more processors to perform the methods described in further detail herein"). Claims 7 and 10 are rejected under U.S.C. §103 as being unpatentable over the combination of Woo and Zheng and Gong (“KDCTime: Knowledge Distillation with Calibration on InceptionTime for Time-series Classification”, 2021). Regarding claim 7, the combination of Woo and Zheng teaches The method of claim 6. However, the combination of Woo and Zheng doesn't explicitly teach wherein the temporal loss is determined by summing Kullback-Leibler divergences between the teacher temporal similarities and the student temporal similarities over all temporal position. Gong, in the same field of endeavor, teaches The method of claim 6, wherein the temporal loss is determined by summing Kullback-Leibler divergences between the teacher temporal similarities and the student temporal similarities over all temporal position.(([p. 5] LKD(yh,ˆy,ytτ,ˆyτ)=(1−ε)LCE(yh,ˆy)+ετ2LKL(ytτ,ˆyτ)"). Zheng explicitly computes teacher and student similarity distribution by comparing the query embedding "to all anchor points", converting those similarities into distributions, and then minimizing ([p. 4] "L=KL(pt||ps)"). Gong then supplies the time-series and then explicitly sums the KL divergence over all teacher and student temporal positions to get the KD Loss term). The combination of Woo and Zheng as well as Gong are directed towards self supervised relational representation learning. Therefore, the combination of Woo and Zheng as well as Gong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Woo and Zheng with the teachings of Gong by using the KL loss term in Zheng in the combined loss formula in Gong for temporal classification. Gong explicitly uses an analogous ResNet architecture further reinforcing the obviousness. Gong provides as additional motivation for combination ([p. 2] “KDCTime simultaneously improves the accuracy and reduces the inference time of it with an acceptable training time overhead. As a conclusion, the performance of KDCTime is promising”). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 10, the combination of Woo and Zheng teaches The method of claim 9. However, the combination of Woo and Zheng doesn't explicitly teach, wherein the instance loss is determined by summing Kullback-Leibler divergences between the teacher instance similarities and the student instance similarities over all temporal position. Gong, in the same field of endeavor, teaches The method of claim 9, wherein the instance loss is determined by summing Kullback-Leibler divergences between the teacher instance similarities and the student instance similarities over all temporal position. (([p. 5] KD(yh,ˆy,ytτ,ˆyτ)=(1−ε)LCE(yh,ˆy)+ετ2LKL(ytτ,ˆyτ)"). Zheng explicitly computes teacher and student similarity distribution by comparing the query embedding "to all anchor points", converting those similarities into distributions, and then minimizing ([p. 4] "L=KL(pt||ps)"). Gong then supplies the time-series and then explicitly sums the KL divergence over all teacher and student temporal positions to get the KD Loss term). The combination of Woo and Zheng as well as Gong are directed towards self supervised relational representation learning. Therefore, the combination of Woo and Zheng as well as Gong are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Woo and Zheng with the teachings of Gong by using the KL loss term in Zheng in the combined loss formula in Gong for temporal classification. Gong explicitly uses an analogous ResNet architecture further reinforcing the obviousness. Gong provides as additional motivation for combination ([p. 2] “KDCTime simultaneously improves the accuracy and reduces the inference time of it with an acceptable training time overhead. As a conclusion, the performance of KDCTime is promising”). This motivation for combination also applies to the remaining claims which depend on this combination. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

May 19, 2023
Application Filed
Mar 17, 2026
Non-Final Rejection mailed — §101, §103, §112
Jun 10, 2026
Response Filed
Aug 21, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~1y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month